
- AI in research
- 5 min read
- By George Burchell
- View publications on PubMed
- ORCID
Where I've Drawn the Line with AI
Over the last few years, I’ve spent a lot of time thinking about where AI genuinely helps in research work, and where it starts to cause problems later on. The problems that show up months down the line, when someone asks a reasonable question and you realise you can’t quite answer it with the confidence you expected.
That tension has pushed me to get firmer about where AI should stop, even when it could technically go further.
This is about understanding how confidence actually works in real reviews, and what people are left holding when the work is challenged.
What AI Is Very Good At (And Why That’s Not the Issue)
In systematic literature reviews, there’s no question that AI excels at enforcing structure. It’s extremely good at the repetitive parts of the process such as: running searches, deduplicating records, supporting screening, templating extraction, keeping audit trails tidy, checking for consistency, and surfacing discrepancies that are easy to miss when you’re tired or under pressure.
It’s also very good at preparing decisions including, highlighting borderline cases, flagging conflicts between reviewers, summarising patterns in the evidence, stress-testing logic, and making assumptions visible that might otherwise stay implicit.
None of that feels controversial anymore. In practice, those capabilities genuinely reduce cognitive load. They stop people wasting energy on admin and bookkeeping when their attention is better spent thinking. But none of those things are where the real line is.
The Difference Between Preparing a Decision and Owning One
There are decisions in a review that don’t just need to be made, but need to be owned. Choices around inclusion and exclusion, outcome prioritisation, how heterogeneity is handled, how bias is weighed, and how uncertainty is interpreted all fall into this category. These are not mechanical steps that can be resolved by following a template; they are judgement calls shaped by context, experience, and how the evidence will be used.
Crucially, these decisions do not stay in the moment they are made. They resurface later, sometimes months or years down the line, when the work is scrutinised, cited, or challenged. Questions arise about why a study was included, why an outcome was prioritised, or why certain evidence carried more weight.
That is where the difference between plausible reasoning and real responsibility becomes clear. AI can generate reasoning that sounds convincing, often extremely so, but plausibility is not the same as accountability. When a decision is questioned long after the review is complete, someone needs to be able to stand behind it clearly, without reconstructing their thinking or relying on a vague memory of what the system produced at the time.
Why Replacing Judgement Increases Anxiety, Not Confidence
One pattern I’ve noticed repeatedly is that systems which replace judgement, rather than support it, tend to create a delayed form of stress. In the moment, everything feels smoother. Decisions disappear into the system, progress feels fast, and there’s a sense of relief as fewer choices land on your desk.
That relief doesn’t always last. When those decisions are questioned later, the confidence you expected to feel often isn’t there. Instead, a low-level anxiety creeps in, trying to recall why something was done, wondering whether you would still defend it, or whether you would make the same call again.
The irony is that tools designed to reduce effort can increase mental load over time, not because they made the wrong decision, but because they removed ownership from the person who ultimately has to answer for it.
Accountability Is the Real Constraint
It’s tempting to treat these boundaries as technical limitations, as if the only reason AI shouldn’t go further is because it isn’t “good enough yet”. But that isn’t the real constraint. The real issue is accountability.
In research, confidence is rarely felt at the moment a decision is made. It shows up later, when the work is questioned. Good systems don’t just help you reach an answer; they help you explain how you reached it without having to excavate your own reasoning.
That’s why the design principle I keep returning to is simple: AI should reduce cognitive load, not decision ownership. When a system supports traceable reasoning and makes decisions easy to justify against first principles, people stay calmer. They can stand by the work clearly, without defensiveness or uncertainty.
Calm Comes from Being Able to Trace Your Thinking
The clearest signal I look for now isn’t speed or sophistication, but whether a system leaves someone feeling steady afterwards. Whether they can trace a decision back without effort, explain it without hedging, and trust the work because they recognise their own judgement in it.
Good systems don’t try to remove humans from the uncomfortable parts of research. They support people through them. They make the hard thinking visible rather than hidden, reducing the background noise that quietly erodes confidence over time.
That’s the line I no longer waver on, because people still have to live with these decisions long after the software has finished its job.

About the Author
Connect on LinkedInGeorge Burchell
George Burchell is a specialist in systematic literature reviews and scientific evidence synthesis with significant expertise in integrating advanced AI technologies and automation tools into the research process. With over four years of consulting and practical experience, he has developed and led multiple projects focused on accelerating and refining the workflow for systematic reviews within medical and scientific research.