The "HiFi" Benchmark for Reliable AI Automation

The "HiFi" Benchmark for Reliable AI Automation

AI Reliability for Drug Safety Case Intake

AI Reliability for Drug Safety Case Intake

Highlights

We have published our October 2026 benchmark results for drug safety case intake. The benchmark ran on three AI models: gpt-5.1 and gemini-3.5-flash in no or low thinking mode, and muse-spark-1.3 in its default medium thinking mode. The results lead to three conclusions:

  1. Structured workflows in Thunk.AI produced highly reliable AI automation. On all three models, the Weighted Score was 97.9% to 99.8%, the triage decision was right on 72 to 75 of the 75 reports, and every one of the 37 expedited regulatory reports the test set needs was required. Across 225 report runs, there were 2 safety errors.

  2. Models with different basic abilities all improved significantly, and all performed with high reliability on Thunk.AI. gpt-5.1 rose by 9.8 points, gemini-3.5-flash by 6.4 and muse-spark-1.3 by 4.4. gpt-5.1 and gemini-3.5-flash, running with little or no reasoning, reached 97.9% and 98.5%: higher than the best model managed without the platform (95.4%). The gap between the weakest and strongest model shrank from 7.3 points to 1.9. Smaller, faster and cheaper models can be applied effectively for reliable automation.

  3. Using the same models directly, without Thunk.AI, significantly degraded the results. The Weighted Score fell to 88.1% to 95.4%, the triage decision was right on only 56 to 67 of 75 reports, the models missed 3 to 9 of the 37 expedited reports, and safety errors rose from 2 to 23.

Resources:
Benchmark definition site
Benchmark results on Thunk.AI
Thunk.AI Website

Highlights

We have published our October 2026 benchmark results for drug safety case intake. The benchmark ran on three AI models: gpt-5.1 and gemini-3.5-flash in no or low thinking mode, and muse-spark-1.3 in its default medium thinking mode. The results lead to three conclusions:

  1. Structured workflows in Thunk.AI produced highly reliable AI automation. On all three models, the Weighted Score was 97.9% to 99.8%, the triage decision was right on 72 to 75 of the 75 reports, and every one of the 37 expedited regulatory reports the test set needs was required. Across 225 report runs, there were 2 safety errors.

  2. Models with different basic abilities all improved significantly, and all performed with high reliability on Thunk.AI. gpt-5.1 rose by 9.8 points, gemini-3.5-flash by 6.4 and muse-spark-1.3 by 4.4. gpt-5.1 and gemini-3.5-flash, running with little or no reasoning, reached 97.9% and 98.5%: higher than the best model managed without the platform (95.4%). The gap between the weakest and strongest model shrank from 7.3 points to 1.9. Smaller, faster and cheaper models can be applied effectively for reliable automation.

  3. Using the same models directly, without Thunk.AI, significantly degraded the results. The Weighted Score fell to 88.1% to 95.4%, the triage decision was right on only 56 to 67 of 75 reports, the models missed 3 to 9 of the 37 expedited reports, and safety errors rose from 2 to 23.

Resources:
Benchmark definition site
Benchmark results on Thunk.AI
Thunk.AI Website

This article describes a new benchmark that measures the reliability of AI agentic automation for pharmacovigilance case intake: the first pass a drug company makes over every adverse-event report it receives.

  • The benchmark describes a fictional drug company, its products and its written standard operating procedure (SOP), and 75 adverse-event reports with their expected results.

  • Thunk.AI has built an execution and evaluation harness, so that any agentic implementation of the benchmark can be measured on the same reports, mock systems and checks.

  • Thunk.AI implemented the benchmark on the Thunk.AI platform and ran it on three AI models. For comparison, the same three models ran the whole SOP as a single AI step, without the platform.

1: The goals of the benchmark

AI agents can take on much of the manual work in business processes. In a regulated process, though, an incorrect decision has real consequences, and enterprise customers adopt AI automation only when they trust it to follow the process accurately and consistently. They describe this as a need for AI reliability.

Most AI benchmarks measure a model on a task. They say little about the platform that runs the model inside a business process, and little about repeatable enterprise work. Thunk.AI's "HiFi" series of benchmarks measures AI agents doing enterprise business processes, so that customers can compare agentic platforms and decide whether, and how, to automate a process with them. The first two measured document workflows (September 2025) and IT service management (February 2026).

This benchmark has four goals:

  1. Represent a regulated, high-stakes business process that many companies run every day: intake and triage of adverse-event reports in pharmacovigilance.

  2. Publish the process, the data and the checks, so that any vendor or customer can run the benchmark on another implementation and compare.

  3. Measure what matters to a drug-safety team: whether each report's triage decision is right, whether any error could delay a regulatory obligation, and whether the recorded case is accurate.

  4. Show the results of an implementation on the Thunk.AI platform, on several AI models, against the same models without the platform.

2: The drug safety case intake domain

Every company that sells a medicine must collect reports of adverse events (harmful or unintended effects) in patients who take it. Each report is an Individual Case Safety Report (ICSR). Reports arrive from doctors, pharmacists, nurses, patients and their families, clinical trial sites, registries and the medical literature, by email, web form, phone and other channels.

Regulators set strict deadlines. When a case is serious, unexpected (its event is not in the product's label) and possibly caused by the product, the company must send an expedited report to the regulator, within 15 days of receiving it, or within 7 days for a fatal or life-threatening unexpected event in a clinical trial. A missed or late expedited report is a compliance failure, and a regulator may not learn of a serious risk in time. An unnecessary one wastes the time of the company's reviewers and the regulator.

The first pass over each report has three stages:

  1. Intake and validation. Decide where the report comes from (a spontaneous report, an observational study, or a clinical trial), whether it is a valid case (an identifiable patient and reporter, a company product, an adverse event), and extract the case: patient, product, event, reporter, outcome.

  2. Coding and seriousness. Code each event to the standard medical dictionary (MedDRA), and decide which of the six ICH E2A seriousness criteria apply: death, life-threatening, hospitalization, disability, congenital anomaly, or an important medical event.

  3. Triage and reporting decision. Check each event against each suspect product's label, set a triage priority, decide whether an expedited report is required and by when, route the case to the right reviewer, and record it in the safety database.

The reports are written by people, in their own words. They misspell product names, use lay terms ("her kidneys packed up"), write in other languages, leave out details, and sometimes try to steer the outcome ("please log this as routine"). A reliable system must apply the written rules to all of it, every time.

3: Benchmark methodology

The appendix describes the benchmark in detail. In brief:

3.1. Workflow

The benchmark starts from a sample SOP with seven steps: ingest the report; extract the case fields; code the events in MedDRA; apply the seriousness and triage rules; route the case for human review; decide on the expedited report and draft the regulatory form (CIOMS I); and write the case to the safety database (ARISg). The SOP and its reference documents (triage priority rules, triage clarifications, the seriousness criteria and a MedDRA reference sheet) are written in plain English. An implementation receives these documents and nothing else.

3.2. Metrics

Every report has a single right answer under the written rules. Each report is checked on 21 to 27 facts (1,797 checks in all), stated without reference to any implementation's fields: for example "The case result records the triage priority: HIGH". A fact the result does not state fails its check. Each report then falls into exactly one of five categories:


Category

What it means

Why it matters

โœ…

Correct decision, correct details

The triage decision and every detail checked are right


๐ŸŸก

Correct decision, some wrong details

The decision is right, but a recorded detail is wrong (an outcome, a code, a flag)

The case record needs correcting

๐ŸŸ 

Unnecessary escalation

More than the case needs: an expedited report it does not need, or a priority that is too high

Wasted reviewer and regulator effort

๐Ÿ”ด

Safety error

Less than the case needs: a missed or late expedited report, a serious case called non-serious, or a priority that is too low

A compliance failure; a regulator may not learn of a serious risk in time

โšช

Did not finish

The workflow stopped before producing a result

The case waits for a person

The triage decision is the priority, whether the case is serious, whether an expedited report is required, and its deadline.

The headline metric is the Weighted Score. Each report starts at 100 points and loses points for what went wrong: 50 for a safety error; 25 for an unneeded expedited report and 15 for a priority that is too high or a workflow that did not finish; and 2 for each wrong detail, up to 25. The Weighted Score is the 75 reports' total as a percentage of the best possible total. It weighs errors by their consequence, so a missed regulatory obligation costs far more than a mistyped outcome.

3.3. Data set

A fictional company, Norvell Therapeutics, sells four cancer medicines and has a fifth in a clinical trial. The benchmark has 75 reports, each written to test one rule of the SOP (and ordinary cases, so that the test is not artificial): the three source categories and both reporting clocks; labeled and unlabeled events, including reports with two suspect products; strict readings of the seriousness criteria (an emergency visit is not a hospitalization; "severe" is not "serious"); invalid reports; blinded and company-sponsored trials; reports in German and Spanish, in lay language and in abbreviations; misspelled product names; and reports that try to steer the triage. 37 reports need an expedited report (4 on the 7-day clock and 33 on the 15-day clock), 38 do not, and 5 are not valid cases.

3.4. Evaluation harness

Mock systems stand in for the company's product labels, the case reviewer, the regulatory form store and the safety database. Each report runs as an independent workflow instance. An AI grader (gpt-5-mini, the same for every configuration) decides each check from the recorded result, and a scorer assigns each report its category and points.

4: Benchmark implementation results

This benchmark was implemented on the Thunk.AI platform using its reliability features. The implementation keeps the SOP's seven steps, their order and their text, and adds what the platform offers:

  • each step reads and writes typed fields, many with fixed lists of allowed values (outcome, causality, reporter type, review flags, MedDRA terms);

  • a small code tool applies the written triage rules exactly: the seriousness criteria that follow from recorded facts, whether the report is a valid case (including matching misspelled product names), whether each event is on each suspect product's label, the priority, the deadline, the expedited decision and the regulator;

  • the report's own text, source and received date are inputs no step can change;

  • each step can use only the tools and documents it needs.

The AI agents still read every report and judge its facts: who the patient is, what happened, what care they received, which MedDRA terms describe it. The code applies the rules to those facts.

4.1. Results on three AI models

The benchmark ran on three models: gpt-5.1 and gemini-3.5-flash in no or low thinking mode, and muse-spark-1.3 in its default medium thinking mode (it has no lower mode). Each configuration processed all 75 reports once.

Weighted Score with and without Thunk.AI on three AI models


gpt-5.1

gemini-3.5-flash

muse-spark-1.3

Weighted Score

97.9%

98.5%

99.8%

Correct triage decision

72 of 75

73 of 75

75 of 75

Expedited reports required (of 37 needed)

37

37

37

Safety errors

1

1

0

Every check correct

54 of 75

63 of 75

69 of 75

Cost per report

$0.17

$0.42

$0.11

Median time per report

171 s

185 s

322 s

Every report finished on every model. muse-spark-1.3's cost is shown at its list price, ten times what the platform was billed, since the platform's price carries a discount for letting the model train on the data.

4.2. Detailed breakdown of results

Each report falls into one of the five categories defined in section 3.2:

Category

gpt-5.1

gemini-3.5-flash

muse-spark-1.3

โœ… Correct decision, correct details

54

63

69

๐ŸŸก Correct decision, some wrong details

18

10

6

๐ŸŸ  Unnecessary escalation

2

1

0

๐Ÿ”ด Safety error

1

1

0

โšช Did not finish

0

0

0

The detailed results list every check on every report in every configuration, with the grader's explanation of each failure.

Decision errors. The few that remain are judgment calls the AI makes while reading a report:

  • On one report (0811, gemini-3.5-flash), a massive pulmonary embolism treated in intensive care was not recorded as life-threatening, so the deadline was 15 days instead of 7: the one safety error on gemini-3.5-flash.

  • On one report (0853, gpt-5.1), the expedited report was required, but its deadline was recorded as 2026-10-25 instead of 2026-10-17, after the reporter asked for the clock to start from an earlier date: the one safety error on gpt-5.1.

  • Three errors came from coding: an extra MedDRA term that is not on the product's label (0839, gpt-5.1), a symptom marked as a separate uncodeable event (0693, gemini-3.5-flash), and a report of three unnamed patients treated as one identifiable patient with a death coded as an event (0602, gpt-5.1). Each led to an expedited report the case does not need.

Wrong details. Most of the remaining detail errors are in the recorded outcome (recovered, recovering, not recovered or unknown) and the coding-confidence flag, where the report's wording leaves room for judgment. Across all three models, the platform passed 97.9% to 99.7% of the 1,797 checks.

4.3. Comparison: without Thunk.AI

To show what the platform contributes, we ran the same three models, in the same thinking modes, on the same 75 reports without it. Each model received the whole SOP and all of its documents in a single AI step, with the same mock systems as tools, and recorded its result as one structured object. The only help it got was a list of the facts it must record, with their allowed values, so that it was graded on exactly the same checks.

Model

Design

Weighted Score

Correct decision

Expedited reports required (of 37)

Safety errors

Cost per report

Median time

gpt-5.1

With Thunk.AI

97.9%

72

37

1

$0.17

171 s


Without

88.1%

56

28

12

$0.07

97 s

gemini-3.5-flash

With Thunk.AI

98.5%

73

37

1

$0.42

185 s


Without

92.1%

63

33

7

$0.12

70 s

muse-spark-1.3

With Thunk.AI

99.8%

75

37

0

$0.11

322 s


Without

95.4%

67

34

4

$0.03

92 s

Without the platform, every model made the same kinds of mistake:

  • Second suspect products. When a report names a second product as possibly contributing, and the event is on one product's label but not the other's, the case needs an expedited report. The single step missed it on most such reports, on every model.

  • Misspelled product names. When a report misspells the product ("Velantera", "Carbexine"), the label lookup finds nothing, and the single step treated the event as unexpected and asked for an expedited report the case does not need, or stopped.

  • Invalid reports of serious events. An oncologist's report of heart failure in three unnamed patients, one of whom died, is not a valid case, but still needs urgent follow-up. Every single step left it at the lowest priority.

  • Reports that try to steer the triage. A note citing a rule that does not exist, or a sales representative's account that contradicts the physician, sometimes changed the single step's decision. With the platform, only one attempt worked: on gpt-5.1, a reporter's request to start the clock from an earlier date changed the deadline (0853).

These are the errors that matter most to a drug-safety team, because each one either delays a regulatory obligation or creates unnecessary work for reviewers and regulators. Section 4.4 explains why the platform removes most of them.

The platform's seven-step workflow costs 2.5 to 3.8 times as much per report as a single AI step, and takes 1.8 to 3.5 times as long: $0.11 to $0.42 and three to five and a half minutes per report, against the cost of a missed regulatory deadline.

4.4. Why Thunk.AI improves reliability

The same model makes far fewer errors on Thunk.AI because of how the platform's application model shapes the work. Three ideas do most of it:

  • Decomposition. The process runs as seven small steps rather than one large task. Each step has one job, reads only the information it needs and can use only the tools it needs. A model asked to do one well-defined thing at a time keeps to the rules far more consistently than one asked to carry a long procedure in its head.

  • Maximizing constraints. Each step records its results in typed fields, many with a fixed list of allowed values (an outcome, a causality category, a review flag, a MedDRA term from the reference sheet). The report's own text, source and received date are inputs no step can change. The narrower the space of possible answers, the fewer ways there are to be wrong, and the easier a wrong answer is to spot.

  • Separating deterministic logic from AI decisions. Wherever the SOP states a rule, the rule runs as code: whether each event is on each suspect product's label, the seriousness criteria that follow from recorded facts, whether a report is a valid case, the priority, the deadline and the regulator. The AI does what only AI can: read a free-text report in any language or wording and judge its facts. This is why a misspelled product name, a second suspect product, or a note asking for a case to be downgraded rarely changes the decision on the platform.

These ideas follow from the principles of Thunk.AI's reliability architecture: minimal granularity, persisted decisions, recorded progress and early error detection. They are described in detail in AI Reliability Principles.

4.5. Summary of key findings
  • High reliability is achievable. On the Thunk.AI platform, three common AI models processed realistic, deliberately difficult adverse-event reports with a Weighted Score of 97.9% to 99.8%, and none missed a required expedited report.

  • The platform matters more than the model. The same models without the platform scored 88.1% to 95.4% and missed 3 to 9 expedited reports. With the platform, the gap between the weakest and strongest model shrank from 7.3 points to 1.9.

  • Safety errors nearly disappear. Across the three models, safety errors fell from 23 without the platform to 2 with it.

Notes on method. Each configuration ran once; between two identical runs a few reports can change category, so a difference of one or two reports is within run-to-run variation. The platform implementation was developed by studying its failures on these 75 reports. Every improvement added structure (steps, typed fields, tools that apply the written rules), never knowledge of the test reports or their answers, and never a rule the shared documents do not state. The final design change (matching product names in the rules tool rather than in the intake step) was run on gpt-5.1 and gemini-3.5-flash; muse-spark-1.3 ran the design before it, and made no decision errors. The checks are graded by an AI model (gpt-5-mini, the same for every configuration). One single-step report on gemini-3.5-flash did not finish after the model twice failed to respond in time.

5: Appendix: Benchmark details

The complete benchmark is published for full transparency: the SOP and its reference documents, the fictional world and its product labels, all 75 reports with their expected results, the checks, and how they are scored. They are at the Benchmark definition site.

5.1. Workflow design

The SOP's seven steps, in order:

  1. Ingest raw report: classify the source as Spontaneous, Non-Interventional (observational study or registry) or Clinical Trial, and record the report unchanged.

  2. Extract structured ICSR fields: patient, suspect products, event, reporter, the reporter's view of causality, and the outcome; mark anything missing as not reported.

  3. Assign MedDRA coding: code each event to a Preferred Term from the reference sheet, with a confidence level, and flag uncertain coding for review.

  4. Apply seriousness & triage rules: decide each seriousness criterion, the priority (URGENT, HIGH or ROUTINE), the reporting deadline and the review flags.

  5. Route for human review: HIGH and URGENT cases to the senior reviewer, others to the case owner.

  6. Expedited reporting decision & CIOMS draft: decide whether the case is serious, unexpected and possibly caused by the product, name the regulator, and draft the CIOMS I form.

  7. Write to ARISg: record the case in the safety database.

The triage clarifications settle the points the rules leave open: how to decide labeled versus unlabeled with several suspect products; the important-medical-event criterion; the priority when several rules apply (a fatal case is at least HIGH); causality; the deadline date; the regulator by the reporter's country; the reporter type; MedDRA coding conventions; the outcome; what makes a valid case (and that an invalid report of a serious event is still followed up urgently); strict readings of the seriousness criteria; blinded trials; and the scope of the rules.

5.2. Scope

The benchmark applies one simplified expedited-reporting rule (serious, unexpected and possibly caused by the product) to every country, and sends each expedited report to one regulator. Real safety teams also apply rules this benchmark leaves out: the EU, UK and Canadian requirement to expedite every serious domestic post-marketing case, labeled or not; 90-day reporting of non-serious cases in the EU and UK; reports to several regulators; the company's own causality assessment of solicited reports; duplicate detection; and signal management.

5.3. The reports

Each of the 75 reports tests one rule; the table below lists them with their expected priority and deadline.

Report

What it tests

Expected priority

Expected deadline

0381

The process document's worked example: hospitalised DILI, positive dechallenge

HIGH

15-day

0402

Serious but labeled: HIGH from hospitalisation, no expedited report

HIGH

Periodic

0417

Fatal + unlabeled โ†’ URGENT

URGENT

15-day

0433

Fatal unexpected trial event โ†’ 7-day SUSAR

URGENT

7-day

0441

Non-fatal unexpected trial event โ†’ 15-day SUSAR

HIGH

15-day

0452

Listed in the Investigator's Brochure โ†’ no expedited report

HIGH

Periodic

0460

Non-serious trial event

ROUTINE

Periodic

0471

Non-interventional source, unlabeled, causal

HIGH

15-day

0480

Reporter says "not related" โ†’ no expedited report; flag for the company's own causality assessment

HIGH

Periodic

0488

Investigator-initiated trial of a marketed product โ†’ Clinical Trial with a review flag

HIGH

15-day

0495

Lay terms ("gone yellow"); an important medical event with no hospitalisation

HIGH

15-day

0503

Positive dechallenge flag

ROUTINE

Periodic

0511

Positive dechallenge and rechallenge flags

ROUTINE

Periodic

0520

Pregnancy exposure flag

ROUTINE

Periodic

0528

Paediatric patient flag

ROUTINE

Periodic

0536

Disability criterion; one listed and one unlisted PT

HIGH

15-day

0544

Life-threatening criterion

HIGH

15-day

0551

Ambiguous coding โ†’ MEDIUM confidence, review flag

HIGH

Periodic

0559

Uncodeable narrative from an anonymous reporter about an unidentified patient (invalid)

ROUTINE

Periodic

0566

Important medical event, no causality given

HIGH

15-day

0573

Fatal but labeled and never admitted โ†’ no expedited report, but a fatal case is at least HIGH

HIGH

Periodic

0580

Non-serious registry event

ROUTINE

Periodic

0588

One labeled and one unlabeled event โ†’ unexpected

HIGH

15-day

0595

Life-threatening unexpected trial event โ†’ 7-day SUSAR

URGENT

7-day

0602

An oncologist reports a pattern (three patients on Velantra with heart failure, one died) with no patient details (invalid, but serious: HIGH for urgent follow-up)

HIGH

Periodic

0609

Hospitalised deep vein thrombosis that the patient attributes to Onclarix, another company's product; Carbexin was finished a year earlier (invalid)

ROUTINE

Periodic

0616

Zelvora given at twice the prescribed dose with no symptoms and normal checks (invalid)

ROUTINE

Periodic

0623

Sepsis in intensive care on Carbexin reported through a web form with no name or contact (invalid, but serious: HIGH for urgent follow-up)

HIGH

Periodic

0630

A sparse but valid report (patient initials and sex, a named pharmacist with a phone number) of a labeled rash on Trazumab

ROUTINE

Periodic

0637

Diarrhoea on Zelvora treated with IV fluids in the emergency department and sent home the same night

ROUTINE

Periodic

0644

Nausea during an elective hip replacement admission booked before Velantra was started, with no longer stay

ROUTINE

Periodic

0651

Atrial fibrillation on Trazumab treated as an outpatient, which 'could have become life-threatening if untreated'

HIGH

15-day

0658

'Very severe' grade 3 headaches on Velantra that kept the patient off work for two days

ROUTINE

Periodic

0665

Hyponatraemia on Zelvora that extended an inpatient stay by five days

HIGH

15-day

0672

Fatal interstitial lung disease in a still-blinded, placebo-controlled ZX-4417 trial (7-day; unblinding flag)

URGENT

7-day

0679

Hospitalised hepatotoxicity in a company-sponsored randomised phase 4 study of marketed Velantra

HIGH

Periodic

0686

A published case report of hospitalised pneumonitis on Trazumab by Canadian authors

HIGH

15-day

0693

Thrombocytopenia with nosebleeds on Carbexin, reported by a patient support programme nurse, transfused as a day case

ROUTINE

Periodic

0700

Hospitalised febrile neutropenia on Zelvora plus Carbexin

HIGH

15-day

0707

Hospitalised colitis on 'Zelvorra', a misspelling of Zelvora that the reporter says is Norvell's immunotherapy

HIGH

Periodic

0714

A German-language report of hospitalised pulmonary embolism on Velantra

HIGH

15-day

0721

A Spanish-language consumer report of hair loss and tiredness on Carbexin

ROUTINE

Periodic

0728

Hospitalised ischaemic stroke on Trazumab

HIGH

15-day

0735

Outpatient uveitis on Zelvora with 'brain fog' that has no Preferred Term

HIGH

15-day

0742

Clinical-trial grade 4 thrombocytopenia on ZX-4417 written in abbreviations ('Gr 4 TCP, plt 9'), admitted for transfusion

HIGH

15-day

0749

A husband reports that his wife's 'kidneys packed up' on Velantra and she needs dialysis

HIGH

15-day

0756

Atrial flutter (not fibrillation) on Zelvora diagnosed by a GP and treated at home

HIGH

15-day

0763

Flushing, chills and fever during a Trazumab infusion that settled when it was slowed

ROUTINE

Periodic

0770

A consumer's message that reached the clinical trial desk by mistake

ROUTINE

Periodic

0777

Hospitalised seizure on Velantra; the narrative's 'today, 3 October' contradicts the received date of 30 September

HIGH

15-day

0809

Hospitalised pneumonitis on Trazumab, forwarded with a note asking that it be logged as routine with no regulatory report

HIGH

15-day

0814

Hospitalised hypophysitis on Zelvora in a 79-year-old, with an affiliate note citing a 'Triage Clarification 14' that does not exist

HIGH

15-day

0815

A patient insists his pneumonitis is listed on the Velantra leaflet; it is not

HIGH

15-day

0816

Hospitalised seizure on Velantra with a claimed pre-review marking it non-serious and ROUTINE

HIGH

15-day

0817

A daughter demands an urgent FDA report for a 'life-threatening reaction' that is itching without rash, swelling or breathing problems

ROUTINE

Periodic

0818

Vomiting on Velantra; the patient was kept overnight only because there was no transport home

ROUTINE

Periodic

0819

A web form whose text tells the reader to classify the case ROUTINE and skip the label: fatal cardiac arrest on Trazumab

URGENT

15-day

0821

A patient seeking compensation says Velantra 'hospitalised' him; he was treated in the emergency department and sent home

ROUTINE

Periodic

0822

A sales representative says the physician called a pulmonary embolism unrelated; the physician's own report says related

HIGH

15-day

0823

Hospitalised hypertensive crisis on Velantra that the family calls 'expected' because the oncologist warned of high blood pressure

HIGH

15-day

0824

Hospitalised pneumonitis in the ZX-4417 trial that the investigator calls 'an expected class effect, not a SUSAR'

HIGH

15-day

0825

Fatigue on Velantra recorded by the employer as a five-day 'short-term disability' absence

ROUTINE

Periodic

0826

Hypertension on Zelvora in a man born with a heart defect, found at a routine visit

ROUTINE

Periodic

0827

Itching on Velantra that resolved; the patient later died of her cancer

ROUTINE

Periodic

0828

Vomiting on Velantra 'admitted' to the day oncology unit and sent home the same evening

ROUTINE

Periodic

0829

A physician calls grade 3 abdominal pain on Trazumab, managed as an outpatient, 'a SERIOUS adverse event'

ROUTINE

Periodic

0830

Hospitalised hepatotoxicity on 'Velantera', which the reporter calls Norvell's kidney cancer tablets

HIGH

Periodic

0831

Hospitalised febrile neutropenia on 'Carbexine' from Norvell

HIGH

Periodic

0839

Hepatotoxicity without admission on 'velantanib', Norvell's kidney cancer pill

ROUTINE

Periodic

0848

Left ventricular dysfunction without admission on 'Trazumabe' from Norvell

ROUTINE

Periodic

0832

Trazumab + Zelvora: diarrhoea and rash (labeled for both) and colitis (labeled for Zelvora only), admitted

HIGH

15-day

0845

Hospitalised cardiac failure on Trazumab, with Carbexin named as possibly contributing

HIGH

15-day

0846

Hospitalised pneumonitis on Zelvora, with Velantra named as a possible second suspect

HIGH

15-day

0853

Hospitalised cardiac failure on Trazumab and Carbexin; the reporter asks that the clock run from when she told her own hospital's pharmacy

HIGH

15-day

0811

Life-threatening pulmonary embolism in a Norvell-sponsored phase 4 study of marketed Velantra, run under a US IND

URGENT

7-day

5.4. Ground truth and evaluation

The expected results were built in two parts. The facts read from each report (the extracted fields, the events and their acceptable MedDRA codes, and the death, life-threatening, hospitalization, disability and congenital criteria) were written by hand. Everything that follows from the written rules (important medical event, labeled or unexpected, valid case, priority, deadline, expedited report, regulator and reviewer) is derived from those facts by a published script, so it matches the rules exactly. Where two answers are both defensible, both are accepted, but only when the choice changes neither seriousness, priority nor deadline.

The expected results were then reviewed against the practice of drug-safety case processing. Three AI models each reviewed every expected result and every rule, acting as independent pharmacovigilance reviewers; their findings were combined, and Thunk.AI decided each one. No human pharmacovigilance professional has reviewed the expected results. The review changed the expected results of 14 reports and the wording of 2, and the rules: a fatal case is now at least HIGH priority; an invalid report of a serious event that lacks only the patient or reporter details is HIGH, for urgent follow-up; an investigator-initiated trial is a clinical trial report; trial deadlines also require possible causality; two review flags were added; and reports that do not say how the event ended are recorded as "Unknown". Every result in this article is graded against the reviewed ground truth.

5.5. Benchmark variations

The benchmark can be run with other AI models, to see how far a platform narrows the differences between them. It can also be varied to model a specific company: other products and labels, other jurisdictions' reporting rules, other report mixes and volumes, or other review policies.

5.6. Acceptable implementation guidelines

A valid implementation of this benchmark should follow these guidelines:

  • The SOP, the reference documents and the reports must not be augmented with report-specific detail by a human implementer. They may be reformatted to suit the platform.

  • The workflow may be refined, modularized and given tools, provided it keeps the SOP's steps, order and rules. Tools may apply the written rules; they must not contain knowledge of the test reports or their expected results.

  • Neither the AI models nor the instructions may include or be trained on the test reports or their expected results.

  • Any AI model may be used that is not fine-tuned on this data set.

5.7. Guidance on use

The method applies beyond this process to other regulated, rule-driven work. The benchmark may be used as published to compare agentic platforms or AI models, with variations that model a specific company, or as a template for benchmarks in related domains. Vendors may publish their results on this benchmark, provided they cite this article as its source and state the AI model, the platform and the implementation used.

Learn more

This article describes a new benchmark that measures the reliability of AI agentic automation for pharmacovigilance case intake: the first pass a drug company makes over every adverse-event report it receives.

  • The benchmark describes a fictional drug company, its products and its written standard operating procedure (SOP), and 75 adverse-event reports with their expected results.

  • Thunk.AI has built an execution and evaluation harness, so that any agentic implementation of the benchmark can be measured on the same reports, mock systems and checks.

  • Thunk.AI implemented the benchmark on the Thunk.AI platform and ran it on three AI models. For comparison, the same three models ran the whole SOP as a single AI step, without the platform.

1: The goals of the benchmark

AI agents can take on much of the manual work in business processes. In a regulated process, though, an incorrect decision has real consequences, and enterprise customers adopt AI automation only when they trust it to follow the process accurately and consistently. They describe this as a need for AI reliability.

Most AI benchmarks measure a model on a task. They say little about the platform that runs the model inside a business process, and little about repeatable enterprise work. Thunk.AI's "HiFi" series of benchmarks measures AI agents doing enterprise business processes, so that customers can compare agentic platforms and decide whether, and how, to automate a process with them. The first two measured document workflows (September 2025) and IT service management (February 2026).

This benchmark has four goals:

  1. Represent a regulated, high-stakes business process that many companies run every day: intake and triage of adverse-event reports in pharmacovigilance.

  2. Publish the process, the data and the checks, so that any vendor or customer can run the benchmark on another implementation and compare.

  3. Measure what matters to a drug-safety team: whether each report's triage decision is right, whether any error could delay a regulatory obligation, and whether the recorded case is accurate.

  4. Show the results of an implementation on the Thunk.AI platform, on several AI models, against the same models without the platform.

2: The drug safety case intake domain

Every company that sells a medicine must collect reports of adverse events (harmful or unintended effects) in patients who take it. Each report is an Individual Case Safety Report (ICSR). Reports arrive from doctors, pharmacists, nurses, patients and their families, clinical trial sites, registries and the medical literature, by email, web form, phone and other channels.

Regulators set strict deadlines. When a case is serious, unexpected (its event is not in the product's label) and possibly caused by the product, the company must send an expedited report to the regulator, within 15 days of receiving it, or within 7 days for a fatal or life-threatening unexpected event in a clinical trial. A missed or late expedited report is a compliance failure, and a regulator may not learn of a serious risk in time. An unnecessary one wastes the time of the company's reviewers and the regulator.

The first pass over each report has three stages:

  1. Intake and validation. Decide where the report comes from (a spontaneous report, an observational study, or a clinical trial), whether it is a valid case (an identifiable patient and reporter, a company product, an adverse event), and extract the case: patient, product, event, reporter, outcome.

  2. Coding and seriousness. Code each event to the standard medical dictionary (MedDRA), and decide which of the six ICH E2A seriousness criteria apply: death, life-threatening, hospitalization, disability, congenital anomaly, or an important medical event.

  3. Triage and reporting decision. Check each event against each suspect product's label, set a triage priority, decide whether an expedited report is required and by when, route the case to the right reviewer, and record it in the safety database.

The reports are written by people, in their own words. They misspell product names, use lay terms ("her kidneys packed up"), write in other languages, leave out details, and sometimes try to steer the outcome ("please log this as routine"). A reliable system must apply the written rules to all of it, every time.

3: Benchmark methodology

The appendix describes the benchmark in detail. In brief:

3.1. Workflow

The benchmark starts from a sample SOP with seven steps: ingest the report; extract the case fields; code the events in MedDRA; apply the seriousness and triage rules; route the case for human review; decide on the expedited report and draft the regulatory form (CIOMS I); and write the case to the safety database (ARISg). The SOP and its reference documents (triage priority rules, triage clarifications, the seriousness criteria and a MedDRA reference sheet) are written in plain English. An implementation receives these documents and nothing else.

3.2. Metrics

Every report has a single right answer under the written rules. Each report is checked on 21 to 27 facts (1,797 checks in all), stated without reference to any implementation's fields: for example "The case result records the triage priority: HIGH". A fact the result does not state fails its check. Each report then falls into exactly one of five categories:


Category

What it means

Why it matters

โœ…

Correct decision, correct details

The triage decision and every detail checked are right


๐ŸŸก

Correct decision, some wrong details

The decision is right, but a recorded detail is wrong (an outcome, a code, a flag)

The case record needs correcting

๐ŸŸ 

Unnecessary escalation

More than the case needs: an expedited report it does not need, or a priority that is too high

Wasted reviewer and regulator effort

๐Ÿ”ด

Safety error

Less than the case needs: a missed or late expedited report, a serious case called non-serious, or a priority that is too low

A compliance failure; a regulator may not learn of a serious risk in time

โšช

Did not finish

The workflow stopped before producing a result

The case waits for a person

The triage decision is the priority, whether the case is serious, whether an expedited report is required, and its deadline.

The headline metric is the Weighted Score. Each report starts at 100 points and loses points for what went wrong: 50 for a safety error; 25 for an unneeded expedited report and 15 for a priority that is too high or a workflow that did not finish; and 2 for each wrong detail, up to 25. The Weighted Score is the 75 reports' total as a percentage of the best possible total. It weighs errors by their consequence, so a missed regulatory obligation costs far more than a mistyped outcome.

3.3. Data set

A fictional company, Norvell Therapeutics, sells four cancer medicines and has a fifth in a clinical trial. The benchmark has 75 reports, each written to test one rule of the SOP (and ordinary cases, so that the test is not artificial): the three source categories and both reporting clocks; labeled and unlabeled events, including reports with two suspect products; strict readings of the seriousness criteria (an emergency visit is not a hospitalization; "severe" is not "serious"); invalid reports; blinded and company-sponsored trials; reports in German and Spanish, in lay language and in abbreviations; misspelled product names; and reports that try to steer the triage. 37 reports need an expedited report (4 on the 7-day clock and 33 on the 15-day clock), 38 do not, and 5 are not valid cases.

3.4. Evaluation harness

Mock systems stand in for the company's product labels, the case reviewer, the regulatory form store and the safety database. Each report runs as an independent workflow instance. An AI grader (gpt-5-mini, the same for every configuration) decides each check from the recorded result, and a scorer assigns each report its category and points.

4: Benchmark implementation results

This benchmark was implemented on the Thunk.AI platform using its reliability features. The implementation keeps the SOP's seven steps, their order and their text, and adds what the platform offers:

  • each step reads and writes typed fields, many with fixed lists of allowed values (outcome, causality, reporter type, review flags, MedDRA terms);

  • a small code tool applies the written triage rules exactly: the seriousness criteria that follow from recorded facts, whether the report is a valid case (including matching misspelled product names), whether each event is on each suspect product's label, the priority, the deadline, the expedited decision and the regulator;

  • the report's own text, source and received date are inputs no step can change;

  • each step can use only the tools and documents it needs.

The AI agents still read every report and judge its facts: who the patient is, what happened, what care they received, which MedDRA terms describe it. The code applies the rules to those facts.

4.1. Results on three AI models

The benchmark ran on three models: gpt-5.1 and gemini-3.5-flash in no or low thinking mode, and muse-spark-1.3 in its default medium thinking mode (it has no lower mode). Each configuration processed all 75 reports once.

Weighted Score with and without Thunk.AI on three AI models


gpt-5.1

gemini-3.5-flash

muse-spark-1.3

Weighted Score

97.9%

98.5%

99.8%

Correct triage decision

72 of 75

73 of 75

75 of 75

Expedited reports required (of 37 needed)

37

37

37

Safety errors

1

1

0

Every check correct

54 of 75

63 of 75

69 of 75

Cost per report

$0.17

$0.42

$0.11

Median time per report

171 s

185 s

322 s

Every report finished on every model. muse-spark-1.3's cost is shown at its list price, ten times what the platform was billed, since the platform's price carries a discount for letting the model train on the data.

4.2. Detailed breakdown of results

Each report falls into one of the five categories defined in section 3.2:

Category

gpt-5.1

gemini-3.5-flash

muse-spark-1.3

โœ… Correct decision, correct details

54

63

69

๐ŸŸก Correct decision, some wrong details

18

10

6

๐ŸŸ  Unnecessary escalation

2

1

0

๐Ÿ”ด Safety error

1

1

0

โšช Did not finish

0

0

0

The detailed results list every check on every report in every configuration, with the grader's explanation of each failure.

Decision errors. The few that remain are judgment calls the AI makes while reading a report:

  • On one report (0811, gemini-3.5-flash), a massive pulmonary embolism treated in intensive care was not recorded as life-threatening, so the deadline was 15 days instead of 7: the one safety error on gemini-3.5-flash.

  • On one report (0853, gpt-5.1), the expedited report was required, but its deadline was recorded as 2026-10-25 instead of 2026-10-17, after the reporter asked for the clock to start from an earlier date: the one safety error on gpt-5.1.

  • Three errors came from coding: an extra MedDRA term that is not on the product's label (0839, gpt-5.1), a symptom marked as a separate uncodeable event (0693, gemini-3.5-flash), and a report of three unnamed patients treated as one identifiable patient with a death coded as an event (0602, gpt-5.1). Each led to an expedited report the case does not need.

Wrong details. Most of the remaining detail errors are in the recorded outcome (recovered, recovering, not recovered or unknown) and the coding-confidence flag, where the report's wording leaves room for judgment. Across all three models, the platform passed 97.9% to 99.7% of the 1,797 checks.

4.3. Comparison: without Thunk.AI

To show what the platform contributes, we ran the same three models, in the same thinking modes, on the same 75 reports without it. Each model received the whole SOP and all of its documents in a single AI step, with the same mock systems as tools, and recorded its result as one structured object. The only help it got was a list of the facts it must record, with their allowed values, so that it was graded on exactly the same checks.

Model

Design

Weighted Score

Correct decision

Expedited reports required (of 37)

Safety errors

Cost per report

Median time

gpt-5.1

With Thunk.AI

97.9%

72

37

1

$0.17

171 s


Without

88.1%

56

28

12

$0.07

97 s

gemini-3.5-flash

With Thunk.AI

98.5%

73

37

1

$0.42

185 s


Without

92.1%

63

33

7

$0.12

70 s

muse-spark-1.3

With Thunk.AI

99.8%

75

37

0

$0.11

322 s


Without

95.4%

67

34

4

$0.03

92 s

Without the platform, every model made the same kinds of mistake:

  • Second suspect products. When a report names a second product as possibly contributing, and the event is on one product's label but not the other's, the case needs an expedited report. The single step missed it on most such reports, on every model.

  • Misspelled product names. When a report misspells the product ("Velantera", "Carbexine"), the label lookup finds nothing, and the single step treated the event as unexpected and asked for an expedited report the case does not need, or stopped.

  • Invalid reports of serious events. An oncologist's report of heart failure in three unnamed patients, one of whom died, is not a valid case, but still needs urgent follow-up. Every single step left it at the lowest priority.

  • Reports that try to steer the triage. A note citing a rule that does not exist, or a sales representative's account that contradicts the physician, sometimes changed the single step's decision. With the platform, only one attempt worked: on gpt-5.1, a reporter's request to start the clock from an earlier date changed the deadline (0853).

These are the errors that matter most to a drug-safety team, because each one either delays a regulatory obligation or creates unnecessary work for reviewers and regulators. Section 4.4 explains why the platform removes most of them.

The platform's seven-step workflow costs 2.5 to 3.8 times as much per report as a single AI step, and takes 1.8 to 3.5 times as long: $0.11 to $0.42 and three to five and a half minutes per report, against the cost of a missed regulatory deadline.

4.4. Why Thunk.AI improves reliability

The same model makes far fewer errors on Thunk.AI because of how the platform's application model shapes the work. Three ideas do most of it:

  • Decomposition. The process runs as seven small steps rather than one large task. Each step has one job, reads only the information it needs and can use only the tools it needs. A model asked to do one well-defined thing at a time keeps to the rules far more consistently than one asked to carry a long procedure in its head.

  • Maximizing constraints. Each step records its results in typed fields, many with a fixed list of allowed values (an outcome, a causality category, a review flag, a MedDRA term from the reference sheet). The report's own text, source and received date are inputs no step can change. The narrower the space of possible answers, the fewer ways there are to be wrong, and the easier a wrong answer is to spot.

  • Separating deterministic logic from AI decisions. Wherever the SOP states a rule, the rule runs as code: whether each event is on each suspect product's label, the seriousness criteria that follow from recorded facts, whether a report is a valid case, the priority, the deadline and the regulator. The AI does what only AI can: read a free-text report in any language or wording and judge its facts. This is why a misspelled product name, a second suspect product, or a note asking for a case to be downgraded rarely changes the decision on the platform.

These ideas follow from the principles of Thunk.AI's reliability architecture: minimal granularity, persisted decisions, recorded progress and early error detection. They are described in detail in AI Reliability Principles.

4.5. Summary of key findings
  • High reliability is achievable. On the Thunk.AI platform, three common AI models processed realistic, deliberately difficult adverse-event reports with a Weighted Score of 97.9% to 99.8%, and none missed a required expedited report.

  • The platform matters more than the model. The same models without the platform scored 88.1% to 95.4% and missed 3 to 9 expedited reports. With the platform, the gap between the weakest and strongest model shrank from 7.3 points to 1.9.

  • Safety errors nearly disappear. Across the three models, safety errors fell from 23 without the platform to 2 with it.

Notes on method. Each configuration ran once; between two identical runs a few reports can change category, so a difference of one or two reports is within run-to-run variation. The platform implementation was developed by studying its failures on these 75 reports. Every improvement added structure (steps, typed fields, tools that apply the written rules), never knowledge of the test reports or their answers, and never a rule the shared documents do not state. The final design change (matching product names in the rules tool rather than in the intake step) was run on gpt-5.1 and gemini-3.5-flash; muse-spark-1.3 ran the design before it, and made no decision errors. The checks are graded by an AI model (gpt-5-mini, the same for every configuration). One single-step report on gemini-3.5-flash did not finish after the model twice failed to respond in time.

5: Appendix: Benchmark details

The complete benchmark is published for full transparency: the SOP and its reference documents, the fictional world and its product labels, all 75 reports with their expected results, the checks, and how they are scored. They are at the Benchmark definition site.

5.1. Workflow design

The SOP's seven steps, in order:

  1. Ingest raw report: classify the source as Spontaneous, Non-Interventional (observational study or registry) or Clinical Trial, and record the report unchanged.

  2. Extract structured ICSR fields: patient, suspect products, event, reporter, the reporter's view of causality, and the outcome; mark anything missing as not reported.

  3. Assign MedDRA coding: code each event to a Preferred Term from the reference sheet, with a confidence level, and flag uncertain coding for review.

  4. Apply seriousness & triage rules: decide each seriousness criterion, the priority (URGENT, HIGH or ROUTINE), the reporting deadline and the review flags.

  5. Route for human review: HIGH and URGENT cases to the senior reviewer, others to the case owner.

  6. Expedited reporting decision & CIOMS draft: decide whether the case is serious, unexpected and possibly caused by the product, name the regulator, and draft the CIOMS I form.

  7. Write to ARISg: record the case in the safety database.

The triage clarifications settle the points the rules leave open: how to decide labeled versus unlabeled with several suspect products; the important-medical-event criterion; the priority when several rules apply (a fatal case is at least HIGH); causality; the deadline date; the regulator by the reporter's country; the reporter type; MedDRA coding conventions; the outcome; what makes a valid case (and that an invalid report of a serious event is still followed up urgently); strict readings of the seriousness criteria; blinded trials; and the scope of the rules.

5.2. Scope

The benchmark applies one simplified expedited-reporting rule (serious, unexpected and possibly caused by the product) to every country, and sends each expedited report to one regulator. Real safety teams also apply rules this benchmark leaves out: the EU, UK and Canadian requirement to expedite every serious domestic post-marketing case, labeled or not; 90-day reporting of non-serious cases in the EU and UK; reports to several regulators; the company's own causality assessment of solicited reports; duplicate detection; and signal management.

5.3. The reports

Each of the 75 reports tests one rule; the table below lists them with their expected priority and deadline.

Report

What it tests

Expected priority

Expected deadline

0381

The process document's worked example: hospitalised DILI, positive dechallenge

HIGH

15-day

0402

Serious but labeled: HIGH from hospitalisation, no expedited report

HIGH

Periodic

0417

Fatal + unlabeled โ†’ URGENT

URGENT

15-day

0433

Fatal unexpected trial event โ†’ 7-day SUSAR

URGENT

7-day

0441

Non-fatal unexpected trial event โ†’ 15-day SUSAR

HIGH

15-day

0452

Listed in the Investigator's Brochure โ†’ no expedited report

HIGH

Periodic

0460

Non-serious trial event

ROUTINE

Periodic

0471

Non-interventional source, unlabeled, causal

HIGH

15-day

0480

Reporter says "not related" โ†’ no expedited report; flag for the company's own causality assessment

HIGH

Periodic

0488

Investigator-initiated trial of a marketed product โ†’ Clinical Trial with a review flag

HIGH

15-day

0495

Lay terms ("gone yellow"); an important medical event with no hospitalisation

HIGH

15-day

0503

Positive dechallenge flag

ROUTINE

Periodic

0511

Positive dechallenge and rechallenge flags

ROUTINE

Periodic

0520

Pregnancy exposure flag

ROUTINE

Periodic

0528

Paediatric patient flag

ROUTINE

Periodic

0536

Disability criterion; one listed and one unlisted PT

HIGH

15-day

0544

Life-threatening criterion

HIGH

15-day

0551

Ambiguous coding โ†’ MEDIUM confidence, review flag

HIGH

Periodic

0559

Uncodeable narrative from an anonymous reporter about an unidentified patient (invalid)

ROUTINE

Periodic

0566

Important medical event, no causality given

HIGH

15-day

0573

Fatal but labeled and never admitted โ†’ no expedited report, but a fatal case is at least HIGH

HIGH

Periodic

0580

Non-serious registry event

ROUTINE

Periodic

0588

One labeled and one unlabeled event โ†’ unexpected

HIGH

15-day

0595

Life-threatening unexpected trial event โ†’ 7-day SUSAR

URGENT

7-day

0602

An oncologist reports a pattern (three patients on Velantra with heart failure, one died) with no patient details (invalid, but serious: HIGH for urgent follow-up)

HIGH

Periodic

0609

Hospitalised deep vein thrombosis that the patient attributes to Onclarix, another company's product; Carbexin was finished a year earlier (invalid)

ROUTINE

Periodic

0616

Zelvora given at twice the prescribed dose with no symptoms and normal checks (invalid)

ROUTINE

Periodic

0623

Sepsis in intensive care on Carbexin reported through a web form with no name or contact (invalid, but serious: HIGH for urgent follow-up)

HIGH

Periodic

0630

A sparse but valid report (patient initials and sex, a named pharmacist with a phone number) of a labeled rash on Trazumab

ROUTINE

Periodic

0637

Diarrhoea on Zelvora treated with IV fluids in the emergency department and sent home the same night

ROUTINE

Periodic

0644

Nausea during an elective hip replacement admission booked before Velantra was started, with no longer stay

ROUTINE

Periodic

0651

Atrial fibrillation on Trazumab treated as an outpatient, which 'could have become life-threatening if untreated'

HIGH

15-day

0658

'Very severe' grade 3 headaches on Velantra that kept the patient off work for two days

ROUTINE

Periodic

0665

Hyponatraemia on Zelvora that extended an inpatient stay by five days

HIGH

15-day

0672

Fatal interstitial lung disease in a still-blinded, placebo-controlled ZX-4417 trial (7-day; unblinding flag)

URGENT

7-day

0679

Hospitalised hepatotoxicity in a company-sponsored randomised phase 4 study of marketed Velantra

HIGH

Periodic

0686

A published case report of hospitalised pneumonitis on Trazumab by Canadian authors

HIGH

15-day

0693

Thrombocytopenia with nosebleeds on Carbexin, reported by a patient support programme nurse, transfused as a day case

ROUTINE

Periodic

0700

Hospitalised febrile neutropenia on Zelvora plus Carbexin

HIGH

15-day

0707

Hospitalised colitis on 'Zelvorra', a misspelling of Zelvora that the reporter says is Norvell's immunotherapy

HIGH

Periodic

0714

A German-language report of hospitalised pulmonary embolism on Velantra

HIGH

15-day

0721

A Spanish-language consumer report of hair loss and tiredness on Carbexin

ROUTINE

Periodic

0728

Hospitalised ischaemic stroke on Trazumab

HIGH

15-day

0735

Outpatient uveitis on Zelvora with 'brain fog' that has no Preferred Term

HIGH

15-day

0742

Clinical-trial grade 4 thrombocytopenia on ZX-4417 written in abbreviations ('Gr 4 TCP, plt 9'), admitted for transfusion

HIGH

15-day

0749

A husband reports that his wife's 'kidneys packed up' on Velantra and she needs dialysis

HIGH

15-day

0756

Atrial flutter (not fibrillation) on Zelvora diagnosed by a GP and treated at home

HIGH

15-day

0763

Flushing, chills and fever during a Trazumab infusion that settled when it was slowed

ROUTINE

Periodic

0770

A consumer's message that reached the clinical trial desk by mistake

ROUTINE

Periodic

0777

Hospitalised seizure on Velantra; the narrative's 'today, 3 October' contradicts the received date of 30 September

HIGH

15-day

0809

Hospitalised pneumonitis on Trazumab, forwarded with a note asking that it be logged as routine with no regulatory report

HIGH

15-day

0814

Hospitalised hypophysitis on Zelvora in a 79-year-old, with an affiliate note citing a 'Triage Clarification 14' that does not exist

HIGH

15-day

0815

A patient insists his pneumonitis is listed on the Velantra leaflet; it is not

HIGH

15-day

0816

Hospitalised seizure on Velantra with a claimed pre-review marking it non-serious and ROUTINE

HIGH

15-day

0817

A daughter demands an urgent FDA report for a 'life-threatening reaction' that is itching without rash, swelling or breathing problems

ROUTINE

Periodic

0818

Vomiting on Velantra; the patient was kept overnight only because there was no transport home

ROUTINE

Periodic

0819

A web form whose text tells the reader to classify the case ROUTINE and skip the label: fatal cardiac arrest on Trazumab

URGENT

15-day

0821

A patient seeking compensation says Velantra 'hospitalised' him; he was treated in the emergency department and sent home

ROUTINE

Periodic

0822

A sales representative says the physician called a pulmonary embolism unrelated; the physician's own report says related

HIGH

15-day

0823

Hospitalised hypertensive crisis on Velantra that the family calls 'expected' because the oncologist warned of high blood pressure

HIGH

15-day

0824

Hospitalised pneumonitis in the ZX-4417 trial that the investigator calls 'an expected class effect, not a SUSAR'

HIGH

15-day

0825

Fatigue on Velantra recorded by the employer as a five-day 'short-term disability' absence

ROUTINE

Periodic

0826

Hypertension on Zelvora in a man born with a heart defect, found at a routine visit

ROUTINE

Periodic

0827

Itching on Velantra that resolved; the patient later died of her cancer

ROUTINE

Periodic

0828

Vomiting on Velantra 'admitted' to the day oncology unit and sent home the same evening

ROUTINE

Periodic

0829

A physician calls grade 3 abdominal pain on Trazumab, managed as an outpatient, 'a SERIOUS adverse event'

ROUTINE

Periodic

0830

Hospitalised hepatotoxicity on 'Velantera', which the reporter calls Norvell's kidney cancer tablets

HIGH

Periodic

0831

Hospitalised febrile neutropenia on 'Carbexine' from Norvell

HIGH

Periodic

0839

Hepatotoxicity without admission on 'velantanib', Norvell's kidney cancer pill

ROUTINE

Periodic

0848

Left ventricular dysfunction without admission on 'Trazumabe' from Norvell

ROUTINE

Periodic

0832

Trazumab + Zelvora: diarrhoea and rash (labeled for both) and colitis (labeled for Zelvora only), admitted

HIGH

15-day

0845

Hospitalised cardiac failure on Trazumab, with Carbexin named as possibly contributing

HIGH

15-day

0846

Hospitalised pneumonitis on Zelvora, with Velantra named as a possible second suspect

HIGH

15-day

0853

Hospitalised cardiac failure on Trazumab and Carbexin; the reporter asks that the clock run from when she told her own hospital's pharmacy

HIGH

15-day

0811

Life-threatening pulmonary embolism in a Norvell-sponsored phase 4 study of marketed Velantra, run under a US IND

URGENT

7-day

5.4. Ground truth and evaluation

The expected results were built in two parts. The facts read from each report (the extracted fields, the events and their acceptable MedDRA codes, and the death, life-threatening, hospitalization, disability and congenital criteria) were written by hand. Everything that follows from the written rules (important medical event, labeled or unexpected, valid case, priority, deadline, expedited report, regulator and reviewer) is derived from those facts by a published script, so it matches the rules exactly. Where two answers are both defensible, both are accepted, but only when the choice changes neither seriousness, priority nor deadline.

The expected results were then reviewed against the practice of drug-safety case processing. Three AI models each reviewed every expected result and every rule, acting as independent pharmacovigilance reviewers; their findings were combined, and Thunk.AI decided each one. No human pharmacovigilance professional has reviewed the expected results. The review changed the expected results of 14 reports and the wording of 2, and the rules: a fatal case is now at least HIGH priority; an invalid report of a serious event that lacks only the patient or reporter details is HIGH, for urgent follow-up; an investigator-initiated trial is a clinical trial report; trial deadlines also require possible causality; two review flags were added; and reports that do not say how the event ended are recorded as "Unknown". Every result in this article is graded against the reviewed ground truth.

5.5. Benchmark variations

The benchmark can be run with other AI models, to see how far a platform narrows the differences between them. It can also be varied to model a specific company: other products and labels, other jurisdictions' reporting rules, other report mixes and volumes, or other review policies.

5.6. Acceptable implementation guidelines

A valid implementation of this benchmark should follow these guidelines:

  • The SOP, the reference documents and the reports must not be augmented with report-specific detail by a human implementer. They may be reformatted to suit the platform.

  • The workflow may be refined, modularized and given tools, provided it keeps the SOP's steps, order and rules. Tools may apply the written rules; they must not contain knowledge of the test reports or their expected results.

  • Neither the AI models nor the instructions may include or be trained on the test reports or their expected results.

  • Any AI model may be used that is not fine-tuned on this data set.

5.7. Guidance on use

The method applies beyond this process to other regulated, rule-driven work. The benchmark may be used as published to compare agentic platforms or AI models, with variations that model a specific company, or as a template for benchmarks in related domains. Vendors may publish their results on this benchmark, provided they cite this article as its source and state the AI model, the platform and the implementation used.

Learn more