TL;DR — Key Takeaways
- AI-generated code is creating real production risk: 80% of surveyed executives said their organizations traced an incident, outage or customer-impacting defect to AI-generated code in the past year.
- AI coding adoption is already substantial: 83% said more than 10% of their production code was generated using AI, while 28% said AI accounts for more than a quarter of production code.
- Testing practices are not keeping pace: Only 6% said AI-generated code is tested with AI tools all the time or often, even as organizations increase their reliance on AI-assisted development.
A survey of 400 business and engineering executives finds 80% have traced a production incident, outage, or customer-impacting defect to code generated by artificial intelligence (AI) tools in the past 12 months.
Conducted by Wakefield Research on behalf of Sauce Labs, a provider of an application testing platform, the survey finds 83% of respondents work for organizations where more than 10% of the code running in production environments has been generated using AI, with 28% now running more than a quarter of the code in their production environments using AI tools.
A full 93% said they also receive reports on that code, but only 38% receive them regularly. A total of 89% also said they disclose to customers when code has been generated using AI, but only 27% said they do so consistently.
A total of 70% said their organizations now spend more than $1 million annually on AI-assisted application development tools and platforms, with 41% spending $5 million or more. A total of 89% described the return on investment (ROI) in AI as either significantly positive (30%) or somewhat positive (59%). More than half (51%) expect fully autonomous testing and software deployment to be possible in either two (27%) or three years (24%), compared to 8% that believe it is achievable now.
More than half (53%) said deploying AI too quickly is the greater risk than falling behind rivals, while conversely, 47% said falling behind is the greater risk. More than 91% also said software quality has either significantly improved (18%) or somewhat improved (73%) thanks to adoption of AI tools.
Overall, only 12% said leadership was greatly concerned about risks that result from deploying AI-generated or AI-tested code in production environments, compared to 58% that are moderately concerned. Only 12% are extremely concerned that safeguards are not effective at preventing AI-generated or AI-tested features that have quality or safety concerns from reaching production, compared to 81% that are very or somewhat concerned.
Additionally, the survey finds a full 84% of respondents lead organizations that have eliminated or significantly reduced roles because of AI adoption, with 53% having cut junior/entry-level developers, while 42% have cut QA testers. Another 34% have reduced the number of technical writers employed, with 24% having cut automation engineers. A total of 42% have also reduced the number of entry-level software engineering positions, compared to 38% that have increased them.
Sauce Labs CEO Prince Kohli said, overall, the survey results suggest that given the number of production incidents traced back to AI code it is apparent that not enough attention is being paid to the need to test code generated using AI tools before it is deployed in production environments. As the volume of code being generated increases, it’s now more a question of how many more incidents will there be before DevOps teams are required to revisit existing workflows, he added. On the plus side, a full 61% said that code is at least sometimes tested using AI tools, but only 6% said they use these tools all the time or often.
Of course, many DevOps teams have not been consistently testing code for years. The survey, for example, finds 53% of respondents admit their organization sometimes (47%) or frequently (6%) has shipped software to production with known or unresolved testing issues in the past year. Two thirds (66%) said their organization has either definitely (21%) or probably (45%) compromised on quality or testing standards to meet release deadlines in the past 12 months. Nearly two thirds (65%) said the approximate financial impact of their most significant software quality incident in the past 12 months exceeded half a million dollars.
Those software defects also resulted in negative media coverage (40%), loss of a major contract/customer (38%), a regulatory inquiry or fine (34%), customer lawsuit/legal action (33%), stock price impact (30%) and an executive departure (27%).
On the plus side, nearly two thirds (64%) report dedicated quality assurance/testing headcount has increased in the past year, with 18% seeing a significant increase.
As always, the proof is in the quality of the proverbial application pudding. While AI has clearly improved productivity, it is also worth remembering that it only takes one or two disgruntled customers to wipe out any of the business benefits that might have been gained.
Frequently Asked Questions
What did the Sauce Labs survey find about AI-generated code?
The survey found that 80% of respondents had traced a production incident, outage or customer-impacting defect to AI-generated code during the previous 12 months.
How widely is AI-generated code being used in production?
A total of 83% said more than 10% of their production code was generated using AI, while 28% said AI-generated code accounted for more than a quarter of their production environments.
Are organizations adequately testing AI-generated code?
The findings suggest testing remains inconsistent. Although 61% said AI-generated code is at least sometimes tested using AI tools, only 6% said they use those tools all the time or often.

