VAQA

14 May 2026

Developer Distrust In AI Output Is A Management Failure

Developers don’t distrust AI output because the models are bad. They distrust it because nobody put a review gate in front of it, and that’s a management problem you can fix this week.

Key takeaways

  • Distrust in AI generated code usually traces back to no mandatory review step
  • Waiting for a smarter model is not a fix, it’s a delay tactic
  • A simple human review gate restores trust faster than any model upgrade
  • Teams without workflow discipline blame the tool instead of the process
  • Operational fixes here are cheap, fast and entirely within a leader’s control

I’ve built and scaled two companies, Obby and Baluu, and now run FirstMotion, an AI Search and GEO agency where AI generated output touches client facing work every single day. Through VAQA I also sit with leadership teams who are wrestling with exactly this problem: engineers who don’t trust what the AI hands them, and a leadership team unsure whether that’s a tooling issue or a people issue.

This piece explains why it’s almost always a workflow issue, and what a proper review gate actually looks like in practice.

The stat everyone quotes, and what it actually measures

Surveys on developer trust in AI generated code get cited constantly, and the framing is almost always the same: developers don’t trust the output, therefore the tools aren’t good enough yet. I think that framing is backwards.

What these surveys are actually picking up is the absence of a defined checkpoint between AI output and production. When there’s no mandatory review step, every engineer has to individually decide, on every single suggestion, whether to trust it or not. That’s exhausting, inconsistent, and it erodes confidence fast, regardless of how good the underlying model is.

Why this looks like a model problem but isn’t

It’s tempting to blame the model because that’s the visible, nameable thing. “GPT got it wrong” or “Claude hallucinated a function” is a much easier story to tell than “our team has no defined process for checking AI output before it ships.”

But think about how engineering teams have always handled uncertain input, whether it comes from a junior engineer, a contractor, or a script someone found online. Nobody expects a junior engineer’s first pull request to go straight to production without review. Nobody would call that a “junior engineer problem.” It’s a process gap, and the fix has always been the same: you put a review gate in front of anything that hasn’t earned unconditional trust yet.

AI output deserves exactly the same treatment, not because it’s inherently worse, but because it’s new enough that most teams haven’t built the muscle yet.

What a proper review gate actually looks like

I’ve helped several portfolio and advisory clients through VAQA rebuild this workflow, and it tends to follow a similar shape regardless of company size.

Stage What happens Who owns it
Generation AI produces code, copy or analysis The tool, unattended
Mandatory review A named human checks logic, edge cases and tone A senior engineer or lead
Sign off Explicit approval recorded, not implied The reviewer, in writing
Ship Output goes live only after sign off The team, not the AI

The critical column is the third one. Most teams that struggle with AI trust have stages one, two and four, but skip explicit sign off. Review happens informally, inconsistently, and nobody can point to a moment where someone said “yes, this is correct.”

Make that moment explicit and trackable, and the distrust problem shrinks dramatically within a few weeks.

Why waiting for a better model doesn’t solve this

I regularly hear leadership teams say some version of “this will get better once the models improve.” Anthropic and OpenAI genuinely do keep shipping stronger models, that part is true. But a smarter model doesn’t remove the need for review, it just changes the error rate you’re reviewing against.

Even a model that’s right 99% of the time will occasionally be wrong in ways that matter, and without a review gate, your team has no reliable way of catching that 1%. Waiting for the model to improve is a delay tactic dressed up as a technical argument, and it lets leadership avoid the more uncomfortable work of fixing the process.

The operational fix, step by step

If you’re running a team that’s hit this problem, here’s the sequence I’d actually walk through with a client.

  1. Identify every point where AI output currently reaches a customer, a codebase or a decision without a named human check
  2. Assign a specific owner to each of those checkpoints, not a team, a person
  3. Define what “approved” looks like in writing, so review isn’t just a vibe check
  4. Track how often output is rejected or edited at that gate, because that data tells you where trust is actually breaking down
  5. Revisit the process quarterly as the team’s comfort with the tools grows

None of this requires new software or a bigger AI budget. It requires someone in leadership deciding that AI output gets the same operational discipline as any other unverified input.

Trust is built through process, not patience

Distrust in AI output isn’t a sign that the technology has stalled. It’s a sign that the workflow around it hasn’t caught up, and workflow is something leadership controls directly.

I’d rather a team ship slightly slower with a clear review gate than fast with no accountability, because the second approach is how trust erodes and never comes back. Fix the process and the tool stops being the scapegoat.

This is squarely the kind of operational fix I work through with clients under Operational Efficiency at VAQA, alongside AI Implementation more broadly. If your team is stuck blaming the model instead of fixing the gate, get in touch and I’ll help you map the checkpoints that actually matter.

Frequently Asked Questions

Isn’t this just extra process that slows teams down?

A well designed review gate adds minutes, not days, and it replaces the slower, invisible cost of engineers second guessing every AI suggestion individually. In my experience it speeds teams up once it’s embedded, because trust stops being negotiated case by case.

How is this different from normal code review?

It isn’t fundamentally different, which is exactly the point. AI output should sit inside your existing review discipline rather than being treated as a special, untrusted category that gets debated separately every time.

Does FirstMotion deal with this kind of AI output trust issue?

Yes, constantly. FirstMotion produces AI assisted content and analysis for B2B SaaS clients, and the same review gate discipline applies there as it does to code, which is part of why I’m confident this pattern generalises beyond engineering teams.

What size of team does this apply to?

Any size. A two person startup needs a lighter version of the same gate, usually just a founder sign off step, while a larger engineering org needs named reviewers per workstream. The principle scales even when the headcount doesn’t.

Can VAQA help set up a review workflow like this?

Yes, this is core to the Operational Efficiency and AI Implementation advisory work I do. I’ll usually start by mapping where AI output currently touches your product or customers unchecked, then design the smallest gate that actually closes that gap. Read more about what I advise on.

Tom Batting is a Forbes 30 Under 30 entrepreneur, founder of Obby and Baluu, and founder of FirstMotion. He advises founders and leadership teams through VAQA on board advisory, growth, go-to-market, AI implementation, operational efficiency, and finance and fundraising.