Diagram contrasting a central orchestrator controlling many agents with a human–AI co-pilot and audit trail

Power vs Trust: What AI Systematic Reviews Get Wrong About Automation

There’s a question I keep coming back to when I think about AI systems: how much control should sit in one place?

In systematic reviews, that question carries real weight. The work isn’t just technical, but also defensible. You’re trying to explain, clearly and calmly, how you got there, long after the work is done. Often to people who are tired, sceptical, or looking for reasons to push back.

That’s why I’ve been spending time thinking about multi-agent systems. Different specialist agents handle different parts of the pipeline, with a single orchestrator sitting above them; one “conductor” keeping the whole review in time.

On paper, it looks like progress, with a single plan and a clear source of truth. However, relying on it feels very different in practice. When control sits in one place, responsibility does too, and a single wrong decision can affect the entire review. In research, that level of risk matters.

Why Coordination Starts to Matter

Anyone who has worked on a complex review will recognise the moment when things slowly begin to drift out of alignment. The search strategy no longer quite matches the protocol, and screening decisions start to reflect a slightly different interpretation of the original question. Over time, even data extraction assumptions can change, often without anyone consciously deciding that they should.

None of this feels dramatic when it happens. It does not announce itself as a problem. Instead, it quietly accumulates in the background, becoming harder to spot the further along you get.

This is where the idea of a central orchestrator starts to look genuinely useful. A system like this can hold the overall plan in mind, remember what was agreed earlier, and notice when later decisions no longer quite fit. It can prevent what I think of as “agent spaghetti”, where lots of activity is happening confidently, but without a shared direction.

Coherence is a powerful thing in this context. When a system remembers its own reasoning, you spend far less time re-orienting yourself. You are not constantly stopping to ask why a decision was made or whether it still makes sense. That shift alone can remove a surprising amount of background mental noise.

However, the moment you centralise that kind of control, you also centralise responsibility.

When One Small Error Travels

In research, small errors rarely remain small for long. A slightly off PICO definition does not affect just a single decision. Instead, it shapes the search strategy, which then shapes the evidence base, which influences the analysis and, ultimately, the conclusions. By the time the problem becomes visible, it is often buried beneath layers of work that all appear reasonable in isolation.

An orchestrator makes this risk more explicit. When a single system coordinates decisions across the pipeline, one poor judgement can cascade smoothly and efficiently all the way downstream. What makes this difficult is that nothing along the way necessarily looks broken.

This changes how you think about reliability, as you stop asking whether an individual agent can perform its task, and start asking what happens if it is wrong and no one notices at the time.

The uncomfortable part is that speed makes this problem worse rather than better. The faster a system moves, the more quickly errors propagate. Power amplifies good decisions, but it amplifies bad ones just as effectively.

At that point, the problem stops being purely technical and becomes behavioural. The real question becomes how the system behaves when it encounters uncertainty. Does it slow down and ask for help, or does it continue confidently because that is what it has been optimised to do?

The Weight of Being Able to Explain Yourself

Systematic reviews are sometimes described as “research you can defend in court”, and that’s not far off. You don’t just need an answer, you need a clear record of what decisions were made, when, and why.

This is where traceability stops being an abstract ideal and starts affecting day-to-day headspace. When the audit trail is weak, you hesitate before moving on. You double-check things you already checked. You keep mental notes because you don’t quite trust the system to remember for you.

A central orchestrator can help here, but only if it’s designed to show its working. Every decision needs a reason attached to it, and every output needs to point back to its sources. Without that, automation increases cognitive load. You end up spending your energy worrying about what you’ll have to justify later.

Letting the System Say “I Don’t Know”

One of the hardest behaviours to design for in any system is uncertainty. Most systems are very good at producing answers, but far fewer are good at recognising when they should not provide one.

In research, saying “I don’t know” is often the most responsible response. It signals a need to pause, to involve a human judgement, or to look more closely at the assumptions sitting underneath a question. However, uncertainty is uncomfortable, and systems tend to smooth over that discomfort by offering confident answers instead.

An orchestrator that never admits uncertainty can feel impressive at first. Over time, though, it quietly trains users to trust it in situations where they should not. Rather than building confidence, this erodes it. You begin to question which decisions were genuinely sound and which were simply delivered with confidence.

Allowing agents to push decisions back to a human is not a failure of automation. It is a deliberate boundary. Those boundaries are what make collaboration between humans and systems feel safe and sustainable.

The Difference Between a Cockpit and a Tool

There is another tension I keep running into, and it has to do with choice. In theory, offering more options should create more flexibility. In practice, it often creates more fatigue.

If an orchestrator surfaces every possible decision, every branching path, and every configurable parameter, it stops feeling like a tool. Instead, it starts to feel like a cockpit, demanding constant attention and interpretation from the person using it.

When people are already tired, clarity matters more than capability. A system that asks fewer questions at better moments often feels calmer and more supportive than one that offers endless control.

This is where a spectrum starts to form in my thinking. At one end of that spectrum is autopilot: fast, smooth, and a little bit magical. It feels reassuring while everything is going well, but when something goes wrong, it can be difficult to know where to look.

At the other end is a co-pilot: slower, more deliberate, and constantly explaining itself. It may be less impressive in a demo, but it is much easier to live with over time.

Where I’m Currently Landing

I am not fully decided, but I find myself leaning towards the co-pilot side, with strong guardrails in place. I am drawn to systems that propose rather than decide, and that carry out work while continuously showing the reasoning behind it. That evidence is there to support confidence. When you can see why a decision was made, you stop second-guessing yourself and move forward with a quieter mind.

For a short product-side take on the same split — AI drafting with inspectable human checkpoints — see Starting AI-assisted systematic reviews.

In evidence synthesis, that quiet matters. Confidence is what allows careful people to keep working without burning out or losing trust in their own judgement. I suspect I will keep thinking this through for some time, because the trade-off between power and trust does not resolve neatly. It forces you to decide what kind of relationship you want people to have with the systems they rely on.

For now, I am paying close attention to how these choices change the way people feel while they work. Not how fast things move, but how steady they feel. That seems like the right place to begin.

George Burchell

About the Author

Connect on LinkedIn

George Burchell

George Burchell is a specialist in systematic literature reviews and scientific evidence synthesis with significant expertise in integrating advanced AI technologies and automation tools into the research process. With over four years of consulting and practical experience, he has developed and led multiple projects focused on accelerating and refining the workflow for systematic reviews within medical and scientific research.