AI-Generated Code Grades High in Reviews but Fails in Production, New Relic Report Finds

  • New Relic's 2026 State of AI Coding report reveals 94% of leaders rate AI-generated code higher than human-authored code in reviews.
  • 78% of respondents report more production incidents after deploying AI-generated code.
  • 82% experienced at least one production failure tied to AI-generated code in the past six months.
  • 67% of technology leaders state AI now generates or refactors 51-75% of their organization’s weekly code output.
  • 96% of technology leaders rate observability as very or extremely important when working with AI-generated code.

The report highlights a growing tension between the perceived quality of AI-generated code during reviews and its actual performance in production environments. This trend underscores the broader industry shift towards AI-driven software development, where the balance between speed and reliability is increasingly critical. The findings suggest that while AI is rapidly becoming a dominant force in coding, its integration into production workflows requires robust observability and verification processes to prevent operational inefficiencies.

Agent Debt
How the accumulation of unvetted architectural logic from AI-generated code will impact long-term software stability and maintenance costs.
Observability
The pace at which engineering teams will adopt upstream observability practices to mitigate risks associated with AI-generated code.
Trust Dynamics
Whether organizations can sustain high levels of trust in AI-generated code while addressing the operational tax it introduces.