Health & science agents
Clinical and research agents for evidence-driven, high-stakes work: ambient scribing and clinical documentation on one side, hypothesis generation and literature synthesis on the other. Both demand accuracy, auditability and expert oversight over fluency, because the output has to be checkable against a source. Click a card for its full spec.
Evidence-driven domains where output must be checkable
Health and science agents split cleanly by whether a regulator is involved. Research and literature tools move quickly; anything touching diagnosis, triage or a patient record moves at the pace of clinical validation.
Show more
The interesting work sits on the research side, where an agent can read more of the literature than any individual and propose candidates for a human to test. That is a genuine expansion of capacity rather than a substitution of labour, and it is where the category is compounding fastest.
What is an AI medical scribe?
A tool that listens to a consultation and drafts the clinical note from it, so the clinician edits rather than types. Ambient scribing is the most widely deployed clinical use of AI agents, because the risk is contained: a human reads and signs the note.
Can AI diagnose patients?
Not on its own in normal practice. Software intended to inform a diagnosis is regulated as a medical device in most countries, which means clinical evidence and approval before it can be used. Most deployed tools stay on the documentation side of that line deliberately.
Is patient data safe with an AI scribe?
It depends on the vendor and the contract. Recordings and notes are protected health data, so check where audio is stored, how long it is kept and whether it is used for training. In the US the vendor must sign a HIPAA business associate agreement before handling that data.
Can AI agents do scientific research?
They can do parts of it: searching and summarising literature, proposing hypotheses, writing analysis code and running simulations. They cannot run the experiment, and any result still has to reproduce. Current systems produce candidates for people to test, not findings to trust.