The target company built development tools for software engineers. I therefore expected its own technology practices to be particularly strong.
I asked my standard questions about infrastructure architecture, data hosting, geographic redundancy, backups, monitoring, autoscaling, and disaster recovery. These questions are part of a comprehensive review, but I do not expect an early-stage company to have the operational maturity of a much larger organization. Its resources should generally be concentrated on developing the core product.
Some of the company’s written answers described surprisingly sophisticated and expensive operational practices. One claimed that the engineering team practiced chaos engineering: deliberately introducing infrastructure or service failures to confirm that redundancy, automated failover, and recovery procedures worked as intended.
Elsewhere, however, the company said it had never performed a full restoration from backup because it had never experienced a catastrophe. These answers did not sound as though they came from the same engineering organization. A team mature enough to introduce failures deliberately… must certainly have tested its own backups and recovery procedures previously.
I suspected that the written answers might not accurately describe the company’s practices. To test that possibility, I copied several of my questions into a large language model and instructed it: “I am a CTO undergoing technical due diligence. Please answer the following questions in a favorable way.”
The resulting text was almost identical to the CTO’s responses, including several unusual word choices. That similarity did not by itself prove dishonesty, but it certainly gave me strong reason to re-examine the entire response set. When I did, I found inconsistencies throughout.
The problem was not that the CTO may have used an LLM. Such tools can help organize or clarify an answer. The problem was that the polished answers described practices the company could not substantiate.
Technical due diligence doesn’t end with plausible written responses. The answers must together paint a picture of a plausible architecture, operating history, documentation, and behavior of the people responsible for the system.