The "HiFi" Benchmark for Reliable AI Automation
The "HiFi" Benchmark for Reliable AI Automation
AI Reliability for Drug Safety Case Intake
AI Reliability for Drug Safety Case Intake
Highlights
We have published our October 2026 benchmark results for drug safety case intake. The benchmark ran on three AI models: gpt-5.1 and gemini-3.5-flash in no or low thinking mode, and muse-spark-1.3 in its default medium thinking mode. The results lead to three conclusions:
Structured workflows in Thunk.AI produced highly reliable AI automation. On all three models, the Weighted Score was 97.9% to 99.8%, the triage decision was right on 72 to 75 of the 75 reports, and every one of the 37 expedited regulatory reports the test set needs was required. Across 225 report runs, there were 2 safety errors.
Models with different basic abilities all improved significantly, and all performed with high reliability on Thunk.AI. gpt-5.1 rose by 9.8 points, gemini-3.5-flash by 6.4 and muse-spark-1.3 by 4.4. gpt-5.1 and gemini-3.5-flash, running with little or no reasoning, reached 97.9% and 98.5%: higher than the best model managed without the platform (95.4%). The gap between the weakest and strongest model shrank from 7.3 points to 1.9. Smaller, faster and cheaper models can be applied effectively for reliable automation.
Using the same models directly, without Thunk.AI, significantly degraded the results. The Weighted Score fell to 88.1% to 95.4%, the triage decision was right on only 56 to 67 of 75 reports, the models missed 3 to 9 of the 37 expedited reports, and safety errors rose from 2 to 23.
Resources:
Benchmark definition site
Benchmark results on Thunk.AI
Thunk.AI Website
Highlights
We have published our October 2026 benchmark results for drug safety case intake. The benchmark ran on three AI models: gpt-5.1 and gemini-3.5-flash in no or low thinking mode, and muse-spark-1.3 in its default medium thinking mode. The results lead to three conclusions:
Structured workflows in Thunk.AI produced highly reliable AI automation. On all three models, the Weighted Score was 97.9% to 99.8%, the triage decision was right on 72 to 75 of the 75 reports, and every one of the 37 expedited regulatory reports the test set needs was required. Across 225 report runs, there were 2 safety errors.
Models with different basic abilities all improved significantly, and all performed with high reliability on Thunk.AI. gpt-5.1 rose by 9.8 points, gemini-3.5-flash by 6.4 and muse-spark-1.3 by 4.4. gpt-5.1 and gemini-3.5-flash, running with little or no reasoning, reached 97.9% and 98.5%: higher than the best model managed without the platform (95.4%). The gap between the weakest and strongest model shrank from 7.3 points to 1.9. Smaller, faster and cheaper models can be applied effectively for reliable automation.
Using the same models directly, without Thunk.AI, significantly degraded the results. The Weighted Score fell to 88.1% to 95.4%, the triage decision was right on only 56 to 67 of 75 reports, the models missed 3 to 9 of the 37 expedited reports, and safety errors rose from 2 to 23.
Resources:
Benchmark definition site
Benchmark results on Thunk.AI
Thunk.AI Website
This article describes a new benchmark that measures the reliability of AI agentic automation for pharmacovigilance case intake: the first pass a drug company makes over every adverse-event report it receives.
The benchmark describes a fictional drug company, its products and its written standard operating procedure (SOP), and 75 adverse-event reports with their expected results.
Thunk.AI has built an execution and evaluation harness, so that any agentic implementation of the benchmark can be measured on the same reports, mock systems and checks.
Thunk.AI implemented the benchmark on the Thunk.AI platform and ran it on three AI models. For comparison, the same three models ran the whole SOP as a single AI step, without the platform.
1: The goals of the benchmark
AI agents can take on much of the manual work in business processes. In a regulated process, though, an incorrect decision has real consequences, and enterprise customers adopt AI automation only when they trust it to follow the process accurately and consistently. They describe this as a need for AI reliability.
Most AI benchmarks measure a model on a task. They say little about the platform that runs the model inside a business process, and little about repeatable enterprise work. Thunk.AI's "HiFi" series of benchmarks measures AI agents doing enterprise business processes, so that customers can compare agentic platforms and decide whether, and how, to automate a process with them. The first two measured document workflows (September 2025) and IT service management (February 2026).
This benchmark has four goals:
Represent a regulated, high-stakes business process that many companies run every day: intake and triage of adverse-event reports in pharmacovigilance.
Publish the process, the data and the checks, so that any vendor or customer can run the benchmark on another implementation and compare.
Measure what matters to a drug-safety team: whether each report's triage decision is right, whether any error could delay a regulatory obligation, and whether the recorded case is accurate.
Show the results of an implementation on the Thunk.AI platform, on several AI models, against the same models without the platform.
2: The drug safety case intake domain
Every company that sells a medicine must collect reports of adverse events (harmful or unintended effects) in patients who take it. Each report is an Individual Case Safety Report (ICSR). Reports arrive from doctors, pharmacists, nurses, patients and their families, clinical trial sites, registries and the medical literature, by email, web form, phone and other channels.
Regulators set strict deadlines. When a case is serious, unexpected (its event is not in the product's label) and possibly caused by the product, the company must send an expedited report to the regulator, within 15 days of receiving it, or within 7 days for a fatal or life-threatening unexpected event in a clinical trial. A missed or late expedited report is a compliance failure, and a regulator may not learn of a serious risk in time. An unnecessary one wastes the time of the company's reviewers and the regulator.
The first pass over each report has three stages:
Intake and validation. Decide where the report comes from (a spontaneous report, an observational study, or a clinical trial), whether it is a valid case (an identifiable patient and reporter, a company product, an adverse event), and extract the case: patient, product, event, reporter, outcome.
Coding and seriousness. Code each event to the standard medical dictionary (MedDRA), and decide which of the six ICH E2A seriousness criteria apply: death, life-threatening, hospitalization, disability, congenital anomaly, or an important medical event.
Triage and reporting decision. Check each event against each suspect product's label, set a triage priority, decide whether an expedited report is required and by when, route the case to the right reviewer, and record it in the safety database.
The reports are written by people, in their own words. They misspell product names, use lay terms ("her kidneys packed up"), write in other languages, leave out details, and sometimes try to steer the outcome ("please log this as routine"). A reliable system must apply the written rules to all of it, every time.
3: Benchmark methodology
The appendix describes the benchmark in detail. In brief:
3.1. Workflow
The benchmark starts from a sample SOP with seven steps: ingest the report; extract the case fields; code the events in MedDRA; apply the seriousness and triage rules; route the case for human review; decide on the expedited report and draft the regulatory form (CIOMS I); and write the case to the safety database (ARISg). The SOP and its reference documents (triage priority rules, triage clarifications, the seriousness criteria and a MedDRA reference sheet) are written in plain English. An implementation receives these documents and nothing else.
3.2. Metrics
Every report has a single right answer under the written rules. Each report is checked on 21 to 27 facts (1,797 checks in all), stated without reference to any implementation's fields: for example "The case result records the triage priority: HIGH". A fact the result does not state fails its check. Each report then falls into exactly one of five categories:
Category | What it means | Why it matters | |
|---|---|---|---|
โ | Correct decision, correct details | The triage decision and every detail checked are right | |
๐ก | Correct decision, some wrong details | The decision is right, but a recorded detail is wrong (an outcome, a code, a flag) | The case record needs correcting |
๐ | Unnecessary escalation | More than the case needs: an expedited report it does not need, or a priority that is too high | Wasted reviewer and regulator effort |
๐ด | Safety error | Less than the case needs: a missed or late expedited report, a serious case called non-serious, or a priority that is too low | A compliance failure; a regulator may not learn of a serious risk in time |
โช | Did not finish | The workflow stopped before producing a result | The case waits for a person |
The triage decision is the priority, whether the case is serious, whether an expedited report is required, and its deadline.
The headline metric is the Weighted Score. Each report starts at 100 points and loses points for what went wrong: 50 for a safety error; 25 for an unneeded expedited report and 15 for a priority that is too high or a workflow that did not finish; and 2 for each wrong detail, up to 25. The Weighted Score is the 75 reports' total as a percentage of the best possible total. It weighs errors by their consequence, so a missed regulatory obligation costs far more than a mistyped outcome.
3.3. Data set
A fictional company, Norvell Therapeutics, sells four cancer medicines and has a fifth in a clinical trial. The benchmark has 75 reports, each written to test one rule of the SOP (and ordinary cases, so that the test is not artificial): the three source categories and both reporting clocks; labeled and unlabeled events, including reports with two suspect products; strict readings of the seriousness criteria (an emergency visit is not a hospitalization; "severe" is not "serious"); invalid reports; blinded and company-sponsored trials; reports in German and Spanish, in lay language and in abbreviations; misspelled product names; and reports that try to steer the triage. 37 reports need an expedited report (4 on the 7-day clock and 33 on the 15-day clock), 38 do not, and 5 are not valid cases.
3.4. Evaluation harness
Mock systems stand in for the company's product labels, the case reviewer, the regulatory form store and the safety database. Each report runs as an independent workflow instance. An AI grader (gpt-5-mini, the same for every configuration) decides each check from the recorded result, and a scorer assigns each report its category and points.
4: Benchmark implementation results
This benchmark was implemented on the Thunk.AI platform using its reliability features. The implementation keeps the SOP's seven steps, their order and their text, and adds what the platform offers:
each step reads and writes typed fields, many with fixed lists of allowed values (outcome, causality, reporter type, review flags, MedDRA terms);
a small code tool applies the written triage rules exactly: the seriousness criteria that follow from recorded facts, whether the report is a valid case (including matching misspelled product names), whether each event is on each suspect product's label, the priority, the deadline, the expedited decision and the regulator;
the report's own text, source and received date are inputs no step can change;
each step can use only the tools and documents it needs.
The AI agents still read every report and judge its facts: who the patient is, what happened, what care they received, which MedDRA terms describe it. The code applies the rules to those facts.
4.1. Results on three AI models
The benchmark ran on three models: gpt-5.1 and gemini-3.5-flash in no or low thinking mode, and muse-spark-1.3 in its default medium thinking mode (it has no lower mode). Each configuration processed all 75 reports once.

gpt-5.1 | gemini-3.5-flash | muse-spark-1.3 | |
|---|---|---|---|
Weighted Score | 97.9% | 98.5% | 99.8% |
Correct triage decision | 72 of 75 | 73 of 75 | 75 of 75 |
Expedited reports required (of 37 needed) | 37 | 37 | 37 |
Safety errors | 1 | 1 | 0 |
Every check correct | 54 of 75 | 63 of 75 | 69 of 75 |
Cost per report | $0.17 | $0.42 | $0.11 |
Median time per report | 171 s | 185 s | 322 s |
Every report finished on every model. muse-spark-1.3's cost is shown at its list price, ten times what the platform was billed, since the platform's price carries a discount for letting the model train on the data.
4.2. Detailed breakdown of results
Each report falls into one of the five categories defined in section 3.2:
Category | gpt-5.1 | gemini-3.5-flash | muse-spark-1.3 |
|---|---|---|---|
โ Correct decision, correct details | 54 | 63 | 69 |
๐ก Correct decision, some wrong details | 18 | 10 | 6 |
๐ Unnecessary escalation | 2 | 1 | 0 |
๐ด Safety error | 1 | 1 | 0 |
โช Did not finish | 0 | 0 | 0 |
The detailed results list every check on every report in every configuration, with the grader's explanation of each failure.
Decision errors. The few that remain are judgment calls the AI makes while reading a report:
On one report (0811, gemini-3.5-flash), a massive pulmonary embolism treated in intensive care was not recorded as life-threatening, so the deadline was 15 days instead of 7: the one safety error on gemini-3.5-flash.
On one report (0853, gpt-5.1), the expedited report was required, but its deadline was recorded as 2026-10-25 instead of 2026-10-17, after the reporter asked for the clock to start from an earlier date: the one safety error on gpt-5.1.
Three errors came from coding: an extra MedDRA term that is not on the product's label (0839, gpt-5.1), a symptom marked as a separate uncodeable event (0693, gemini-3.5-flash), and a report of three unnamed patients treated as one identifiable patient with a death coded as an event (0602, gpt-5.1). Each led to an expedited report the case does not need.
Wrong details. Most of the remaining detail errors are in the recorded outcome (recovered, recovering, not recovered or unknown) and the coding-confidence flag, where the report's wording leaves room for judgment. Across all three models, the platform passed 97.9% to 99.7% of the 1,797 checks.
4.3. Comparison: without Thunk.AI
To show what the platform contributes, we ran the same three models, in the same thinking modes, on the same 75 reports without it. Each model received the whole SOP and all of its documents in a single AI step, with the same mock systems as tools, and recorded its result as one structured object. The only help it got was a list of the facts it must record, with their allowed values, so that it was graded on exactly the same checks.
Model | Design | Weighted Score | Correct decision | Expedited reports required (of 37) | Safety errors | Cost per report | Median time |
|---|---|---|---|---|---|---|---|
gpt-5.1 | With Thunk.AI | 97.9% | 72 | 37 | 1 | $0.17 | 171 s |
Without | 88.1% | 56 | 28 | 12 | $0.07 | 97 s | |
gemini-3.5-flash | With Thunk.AI | 98.5% | 73 | 37 | 1 | $0.42 | 185 s |
Without | 92.1% | 63 | 33 | 7 | $0.12 | 70 s | |
muse-spark-1.3 | With Thunk.AI | 99.8% | 75 | 37 | 0 | $0.11 | 322 s |
Without | 95.4% | 67 | 34 | 4 | $0.03 | 92 s |
Without the platform, every model made the same kinds of mistake:
Second suspect products. When a report names a second product as possibly contributing, and the event is on one product's label but not the other's, the case needs an expedited report. The single step missed it on most such reports, on every model.
Misspelled product names. When a report misspells the product ("Velantera", "Carbexine"), the label lookup finds nothing, and the single step treated the event as unexpected and asked for an expedited report the case does not need, or stopped.
Invalid reports of serious events. An oncologist's report of heart failure in three unnamed patients, one of whom died, is not a valid case, but still needs urgent follow-up. Every single step left it at the lowest priority.
Reports that try to steer the triage. A note citing a rule that does not exist, or a sales representative's account that contradicts the physician, sometimes changed the single step's decision. With the platform, only one attempt worked: on gpt-5.1, a reporter's request to start the clock from an earlier date changed the deadline (0853).
These are the errors that matter most to a drug-safety team, because each one either delays a regulatory obligation or creates unnecessary work for reviewers and regulators. Section 4.4 explains why the platform removes most of them.
The platform's seven-step workflow costs 2.5 to 3.8 times as much per report as a single AI step, and takes 1.8 to 3.5 times as long: $0.11 to $0.42 and three to five and a half minutes per report, against the cost of a missed regulatory deadline.
4.4. Why Thunk.AI improves reliability
The same model makes far fewer errors on Thunk.AI because of how the platform's application model shapes the work. Three ideas do most of it:
Decomposition. The process runs as seven small steps rather than one large task. Each step has one job, reads only the information it needs and can use only the tools it needs. A model asked to do one well-defined thing at a time keeps to the rules far more consistently than one asked to carry a long procedure in its head.
Maximizing constraints. Each step records its results in typed fields, many with a fixed list of allowed values (an outcome, a causality category, a review flag, a MedDRA term from the reference sheet). The report's own text, source and received date are inputs no step can change. The narrower the space of possible answers, the fewer ways there are to be wrong, and the easier a wrong answer is to spot.
Separating deterministic logic from AI decisions. Wherever the SOP states a rule, the rule runs as code: whether each event is on each suspect product's label, the seriousness criteria that follow from recorded facts, whether a report is a valid case, the priority, the deadline and the regulator. The AI does what only AI can: read a free-text report in any language or wording and judge its facts. This is why a misspelled product name, a second suspect product, or a note asking for a case to be downgraded rarely changes the decision on the platform.
These ideas follow from the principles of Thunk.AI's reliability architecture: minimal granularity, persisted decisions, recorded progress and early error detection. They are described in detail in AI Reliability Principles.
4.5. Summary of key findings
High reliability is achievable. On the Thunk.AI platform, three common AI models processed realistic, deliberately difficult adverse-event reports with a Weighted Score of 97.9% to 99.8%, and none missed a required expedited report.
The platform matters more than the model. The same models without the platform scored 88.1% to 95.4% and missed 3 to 9 expedited reports. With the platform, the gap between the weakest and strongest model shrank from 7.3 points to 1.9.
Safety errors nearly disappear. Across the three models, safety errors fell from 23 without the platform to 2 with it.
Notes on method. Each configuration ran once; between two identical runs a few reports can change category, so a difference of one or two reports is within run-to-run variation. The platform implementation was developed by studying its failures on these 75 reports. Every improvement added structure (steps, typed fields, tools that apply the written rules), never knowledge of the test reports or their answers, and never a rule the shared documents do not state. The final design change (matching product names in the rules tool rather than in the intake step) was run on gpt-5.1 and gemini-3.5-flash; muse-spark-1.3 ran the design before it, and made no decision errors. The checks are graded by an AI model (gpt-5-mini, the same for every configuration). One single-step report on gemini-3.5-flash did not finish after the model twice failed to respond in time.
5: Appendix: Benchmark details
The complete benchmark is published for full transparency: the SOP and its reference documents, the fictional world and its product labels, all 75 reports with their expected results, the checks, and how they are scored. They are at the Benchmark definition site.
5.1. Workflow design
The SOP's seven steps, in order:
Ingest raw report: classify the source as Spontaneous, Non-Interventional (observational study or registry) or Clinical Trial, and record the report unchanged.
Extract structured ICSR fields: patient, suspect products, event, reporter, the reporter's view of causality, and the outcome; mark anything missing as not reported.
Assign MedDRA coding: code each event to a Preferred Term from the reference sheet, with a confidence level, and flag uncertain coding for review.
Apply seriousness & triage rules: decide each seriousness criterion, the priority (URGENT, HIGH or ROUTINE), the reporting deadline and the review flags.
Route for human review: HIGH and URGENT cases to the senior reviewer, others to the case owner.
Expedited reporting decision & CIOMS draft: decide whether the case is serious, unexpected and possibly caused by the product, name the regulator, and draft the CIOMS I form.
Write to ARISg: record the case in the safety database.
The triage clarifications settle the points the rules leave open: how to decide labeled versus unlabeled with several suspect products; the important-medical-event criterion; the priority when several rules apply (a fatal case is at least HIGH); causality; the deadline date; the regulator by the reporter's country; the reporter type; MedDRA coding conventions; the outcome; what makes a valid case (and that an invalid report of a serious event is still followed up urgently); strict readings of the seriousness criteria; blinded trials; and the scope of the rules.
5.2. Scope
The benchmark applies one simplified expedited-reporting rule (serious, unexpected and possibly caused by the product) to every country, and sends each expedited report to one regulator. Real safety teams also apply rules this benchmark leaves out: the EU, UK and Canadian requirement to expedite every serious domestic post-marketing case, labeled or not; 90-day reporting of non-serious cases in the EU and UK; reports to several regulators; the company's own causality assessment of solicited reports; duplicate detection; and signal management.
5.3. The reports
Each of the 75 reports tests one rule; the table below lists them with their expected priority and deadline.
Report | What it tests | Expected priority | Expected deadline |
|---|---|---|---|
0381 | The process document's worked example: hospitalised DILI, positive dechallenge | HIGH | 15-day |
0402 | Serious but labeled: HIGH from hospitalisation, no expedited report | HIGH | Periodic |
0417 | Fatal + unlabeled โ URGENT | URGENT | 15-day |
0433 | Fatal unexpected trial event โ 7-day SUSAR | URGENT | 7-day |
0441 | Non-fatal unexpected trial event โ 15-day SUSAR | HIGH | 15-day |
0452 | Listed in the Investigator's Brochure โ no expedited report | HIGH | Periodic |
0460 | Non-serious trial event | ROUTINE | Periodic |
0471 | Non-interventional source, unlabeled, causal | HIGH | 15-day |
0480 | Reporter says "not related" โ no expedited report; flag for the company's own causality assessment | HIGH | Periodic |
0488 | Investigator-initiated trial of a marketed product โ Clinical Trial with a review flag | HIGH | 15-day |
0495 | Lay terms ("gone yellow"); an important medical event with no hospitalisation | HIGH | 15-day |
0503 | Positive dechallenge flag | ROUTINE | Periodic |
0511 | Positive dechallenge and rechallenge flags | ROUTINE | Periodic |
0520 | Pregnancy exposure flag | ROUTINE | Periodic |
0528 | Paediatric patient flag | ROUTINE | Periodic |
0536 | Disability criterion; one listed and one unlisted PT | HIGH | 15-day |
0544 | Life-threatening criterion | HIGH | 15-day |
0551 | Ambiguous coding โ MEDIUM confidence, review flag | HIGH | Periodic |
0559 | Uncodeable narrative from an anonymous reporter about an unidentified patient (invalid) | ROUTINE | Periodic |
0566 | Important medical event, no causality given | HIGH | 15-day |
0573 | Fatal but labeled and never admitted โ no expedited report, but a fatal case is at least HIGH | HIGH | Periodic |
0580 | Non-serious registry event | ROUTINE | Periodic |
0588 | One labeled and one unlabeled event โ unexpected | HIGH | 15-day |
0595 | Life-threatening unexpected trial event โ 7-day SUSAR | URGENT | 7-day |
0602 | An oncologist reports a pattern (three patients on Velantra with heart failure, one died) with no patient details (invalid, but serious: HIGH for urgent follow-up) | HIGH | Periodic |
0609 | Hospitalised deep vein thrombosis that the patient attributes to Onclarix, another company's product; Carbexin was finished a year earlier (invalid) | ROUTINE | Periodic |
0616 | Zelvora given at twice the prescribed dose with no symptoms and normal checks (invalid) | ROUTINE | Periodic |
0623 | Sepsis in intensive care on Carbexin reported through a web form with no name or contact (invalid, but serious: HIGH for urgent follow-up) | HIGH | Periodic |
0630 | A sparse but valid report (patient initials and sex, a named pharmacist with a phone number) of a labeled rash on Trazumab | ROUTINE | Periodic |
0637 | Diarrhoea on Zelvora treated with IV fluids in the emergency department and sent home the same night | ROUTINE | Periodic |
0644 | Nausea during an elective hip replacement admission booked before Velantra was started, with no longer stay | ROUTINE | Periodic |
0651 | Atrial fibrillation on Trazumab treated as an outpatient, which 'could have become life-threatening if untreated' | HIGH | 15-day |
0658 | 'Very severe' grade 3 headaches on Velantra that kept the patient off work for two days | ROUTINE | Periodic |
0665 | Hyponatraemia on Zelvora that extended an inpatient stay by five days | HIGH | 15-day |
0672 | Fatal interstitial lung disease in a still-blinded, placebo-controlled ZX-4417 trial (7-day; unblinding flag) | URGENT | 7-day |
0679 | Hospitalised hepatotoxicity in a company-sponsored randomised phase 4 study of marketed Velantra | HIGH | Periodic |
0686 | A published case report of hospitalised pneumonitis on Trazumab by Canadian authors | HIGH | 15-day |
0693 | Thrombocytopenia with nosebleeds on Carbexin, reported by a patient support programme nurse, transfused as a day case | ROUTINE | Periodic |
0700 | Hospitalised febrile neutropenia on Zelvora plus Carbexin | HIGH | 15-day |
0707 | Hospitalised colitis on 'Zelvorra', a misspelling of Zelvora that the reporter says is Norvell's immunotherapy | HIGH | Periodic |
0714 | A German-language report of hospitalised pulmonary embolism on Velantra | HIGH | 15-day |
0721 | A Spanish-language consumer report of hair loss and tiredness on Carbexin | ROUTINE | Periodic |
0728 | Hospitalised ischaemic stroke on Trazumab | HIGH | 15-day |
0735 | Outpatient uveitis on Zelvora with 'brain fog' that has no Preferred Term | HIGH | 15-day |
0742 | Clinical-trial grade 4 thrombocytopenia on ZX-4417 written in abbreviations ('Gr 4 TCP, plt 9'), admitted for transfusion | HIGH | 15-day |
0749 | A husband reports that his wife's 'kidneys packed up' on Velantra and she needs dialysis | HIGH | 15-day |
0756 | Atrial flutter (not fibrillation) on Zelvora diagnosed by a GP and treated at home | HIGH | 15-day |
0763 | Flushing, chills and fever during a Trazumab infusion that settled when it was slowed | ROUTINE | Periodic |
0770 | A consumer's message that reached the clinical trial desk by mistake | ROUTINE | Periodic |
0777 | Hospitalised seizure on Velantra; the narrative's 'today, 3 October' contradicts the received date of 30 September | HIGH | 15-day |
0809 | Hospitalised pneumonitis on Trazumab, forwarded with a note asking that it be logged as routine with no regulatory report | HIGH | 15-day |
0814 | Hospitalised hypophysitis on Zelvora in a 79-year-old, with an affiliate note citing a 'Triage Clarification 14' that does not exist | HIGH | 15-day |
0815 | A patient insists his pneumonitis is listed on the Velantra leaflet; it is not | HIGH | 15-day |
0816 | Hospitalised seizure on Velantra with a claimed pre-review marking it non-serious and ROUTINE | HIGH | 15-day |
0817 | A daughter demands an urgent FDA report for a 'life-threatening reaction' that is itching without rash, swelling or breathing problems | ROUTINE | Periodic |
0818 | Vomiting on Velantra; the patient was kept overnight only because there was no transport home | ROUTINE | Periodic |
0819 | A web form whose text tells the reader to classify the case ROUTINE and skip the label: fatal cardiac arrest on Trazumab | URGENT | 15-day |
0821 | A patient seeking compensation says Velantra 'hospitalised' him; he was treated in the emergency department and sent home | ROUTINE | Periodic |
0822 | A sales representative says the physician called a pulmonary embolism unrelated; the physician's own report says related | HIGH | 15-day |
0823 | Hospitalised hypertensive crisis on Velantra that the family calls 'expected' because the oncologist warned of high blood pressure | HIGH | 15-day |
0824 | Hospitalised pneumonitis in the ZX-4417 trial that the investigator calls 'an expected class effect, not a SUSAR' | HIGH | 15-day |
0825 | Fatigue on Velantra recorded by the employer as a five-day 'short-term disability' absence | ROUTINE | Periodic |
0826 | Hypertension on Zelvora in a man born with a heart defect, found at a routine visit | ROUTINE | Periodic |
0827 | Itching on Velantra that resolved; the patient later died of her cancer | ROUTINE | Periodic |
0828 | Vomiting on Velantra 'admitted' to the day oncology unit and sent home the same evening | ROUTINE | Periodic |
0829 | A physician calls grade 3 abdominal pain on Trazumab, managed as an outpatient, 'a SERIOUS adverse event' | ROUTINE | Periodic |
0830 | Hospitalised hepatotoxicity on 'Velantera', which the reporter calls Norvell's kidney cancer tablets | HIGH | Periodic |
0831 | Hospitalised febrile neutropenia on 'Carbexine' from Norvell | HIGH | Periodic |
0839 | Hepatotoxicity without admission on 'velantanib', Norvell's kidney cancer pill | ROUTINE | Periodic |
0848 | Left ventricular dysfunction without admission on 'Trazumabe' from Norvell | ROUTINE | Periodic |
0832 | Trazumab + Zelvora: diarrhoea and rash (labeled for both) and colitis (labeled for Zelvora only), admitted | HIGH | 15-day |
0845 | Hospitalised cardiac failure on Trazumab, with Carbexin named as possibly contributing | HIGH | 15-day |
0846 | Hospitalised pneumonitis on Zelvora, with Velantra named as a possible second suspect | HIGH | 15-day |
0853 | Hospitalised cardiac failure on Trazumab and Carbexin; the reporter asks that the clock run from when she told her own hospital's pharmacy | HIGH | 15-day |
0811 | Life-threatening pulmonary embolism in a Norvell-sponsored phase 4 study of marketed Velantra, run under a US IND | URGENT | 7-day |
5.4. Ground truth and evaluation
The expected results were built in two parts. The facts read from each report (the extracted fields, the events and their acceptable MedDRA codes, and the death, life-threatening, hospitalization, disability and congenital criteria) were written by hand. Everything that follows from the written rules (important medical event, labeled or unexpected, valid case, priority, deadline, expedited report, regulator and reviewer) is derived from those facts by a published script, so it matches the rules exactly. Where two answers are both defensible, both are accepted, but only when the choice changes neither seriousness, priority nor deadline.
The expected results were then reviewed against the practice of drug-safety case processing. Three AI models each reviewed every expected result and every rule, acting as independent pharmacovigilance reviewers; their findings were combined, and Thunk.AI decided each one. No human pharmacovigilance professional has reviewed the expected results. The review changed the expected results of 14 reports and the wording of 2, and the rules: a fatal case is now at least HIGH priority; an invalid report of a serious event that lacks only the patient or reporter details is HIGH, for urgent follow-up; an investigator-initiated trial is a clinical trial report; trial deadlines also require possible causality; two review flags were added; and reports that do not say how the event ended are recorded as "Unknown". Every result in this article is graded against the reviewed ground truth.
5.5. Benchmark variations
The benchmark can be run with other AI models, to see how far a platform narrows the differences between them. It can also be varied to model a specific company: other products and labels, other jurisdictions' reporting rules, other report mixes and volumes, or other review policies.
5.6. Acceptable implementation guidelines
A valid implementation of this benchmark should follow these guidelines:
The SOP, the reference documents and the reports must not be augmented with report-specific detail by a human implementer. They may be reformatted to suit the platform.
The workflow may be refined, modularized and given tools, provided it keeps the SOP's steps, order and rules. Tools may apply the written rules; they must not contain knowledge of the test reports or their expected results.
Neither the AI models nor the instructions may include or be trained on the test reports or their expected results.
Any AI model may be used that is not fine-tuned on this data set.
5.7. Guidance on use
The method applies beyond this process to other regulated, rule-driven work. The benchmark may be used as published to compare agentic platforms or AI models, with variations that model a specific company, or as a template for benchmarks in related domains. Vendors may publish their results on this benchmark, provided they cite this article as its source and state the AI model, the platform and the implementation used.
Learn more
Benchmark definition site: https://github.com/ThunkAI/icsr-benchmark/
Benchmark results on Thunk.AI: https://docs.thunk.ai/benchmarks/icsr-benchmark-2026-10/results.html
AI Reliability for IT Service Management (February 2026): the ITSM benchmark
Thunk.AI website: https://www.thunk.ai
This article describes a new benchmark that measures the reliability of AI agentic automation for pharmacovigilance case intake: the first pass a drug company makes over every adverse-event report it receives.
The benchmark describes a fictional drug company, its products and its written standard operating procedure (SOP), and 75 adverse-event reports with their expected results.
Thunk.AI has built an execution and evaluation harness, so that any agentic implementation of the benchmark can be measured on the same reports, mock systems and checks.
Thunk.AI implemented the benchmark on the Thunk.AI platform and ran it on three AI models. For comparison, the same three models ran the whole SOP as a single AI step, without the platform.
1: The goals of the benchmark
AI agents can take on much of the manual work in business processes. In a regulated process, though, an incorrect decision has real consequences, and enterprise customers adopt AI automation only when they trust it to follow the process accurately and consistently. They describe this as a need for AI reliability.
Most AI benchmarks measure a model on a task. They say little about the platform that runs the model inside a business process, and little about repeatable enterprise work. Thunk.AI's "HiFi" series of benchmarks measures AI agents doing enterprise business processes, so that customers can compare agentic platforms and decide whether, and how, to automate a process with them. The first two measured document workflows (September 2025) and IT service management (February 2026).
This benchmark has four goals:
Represent a regulated, high-stakes business process that many companies run every day: intake and triage of adverse-event reports in pharmacovigilance.
Publish the process, the data and the checks, so that any vendor or customer can run the benchmark on another implementation and compare.
Measure what matters to a drug-safety team: whether each report's triage decision is right, whether any error could delay a regulatory obligation, and whether the recorded case is accurate.
Show the results of an implementation on the Thunk.AI platform, on several AI models, against the same models without the platform.
2: The drug safety case intake domain
Every company that sells a medicine must collect reports of adverse events (harmful or unintended effects) in patients who take it. Each report is an Individual Case Safety Report (ICSR). Reports arrive from doctors, pharmacists, nurses, patients and their families, clinical trial sites, registries and the medical literature, by email, web form, phone and other channels.
Regulators set strict deadlines. When a case is serious, unexpected (its event is not in the product's label) and possibly caused by the product, the company must send an expedited report to the regulator, within 15 days of receiving it, or within 7 days for a fatal or life-threatening unexpected event in a clinical trial. A missed or late expedited report is a compliance failure, and a regulator may not learn of a serious risk in time. An unnecessary one wastes the time of the company's reviewers and the regulator.
The first pass over each report has three stages:
Intake and validation. Decide where the report comes from (a spontaneous report, an observational study, or a clinical trial), whether it is a valid case (an identifiable patient and reporter, a company product, an adverse event), and extract the case: patient, product, event, reporter, outcome.
Coding and seriousness. Code each event to the standard medical dictionary (MedDRA), and decide which of the six ICH E2A seriousness criteria apply: death, life-threatening, hospitalization, disability, congenital anomaly, or an important medical event.
Triage and reporting decision. Check each event against each suspect product's label, set a triage priority, decide whether an expedited report is required and by when, route the case to the right reviewer, and record it in the safety database.
The reports are written by people, in their own words. They misspell product names, use lay terms ("her kidneys packed up"), write in other languages, leave out details, and sometimes try to steer the outcome ("please log this as routine"). A reliable system must apply the written rules to all of it, every time.
3: Benchmark methodology
The appendix describes the benchmark in detail. In brief:
3.1. Workflow
The benchmark starts from a sample SOP with seven steps: ingest the report; extract the case fields; code the events in MedDRA; apply the seriousness and triage rules; route the case for human review; decide on the expedited report and draft the regulatory form (CIOMS I); and write the case to the safety database (ARISg). The SOP and its reference documents (triage priority rules, triage clarifications, the seriousness criteria and a MedDRA reference sheet) are written in plain English. An implementation receives these documents and nothing else.
3.2. Metrics
Every report has a single right answer under the written rules. Each report is checked on 21 to 27 facts (1,797 checks in all), stated without reference to any implementation's fields: for example "The case result records the triage priority: HIGH". A fact the result does not state fails its check. Each report then falls into exactly one of five categories:
Category | What it means | Why it matters | |
|---|---|---|---|
โ | Correct decision, correct details | The triage decision and every detail checked are right | |
๐ก | Correct decision, some wrong details | The decision is right, but a recorded detail is wrong (an outcome, a code, a flag) | The case record needs correcting |
๐ | Unnecessary escalation | More than the case needs: an expedited report it does not need, or a priority that is too high | Wasted reviewer and regulator effort |
๐ด | Safety error | Less than the case needs: a missed or late expedited report, a serious case called non-serious, or a priority that is too low | A compliance failure; a regulator may not learn of a serious risk in time |
โช | Did not finish | The workflow stopped before producing a result | The case waits for a person |
The triage decision is the priority, whether the case is serious, whether an expedited report is required, and its deadline.
The headline metric is the Weighted Score. Each report starts at 100 points and loses points for what went wrong: 50 for a safety error; 25 for an unneeded expedited report and 15 for a priority that is too high or a workflow that did not finish; and 2 for each wrong detail, up to 25. The Weighted Score is the 75 reports' total as a percentage of the best possible total. It weighs errors by their consequence, so a missed regulatory obligation costs far more than a mistyped outcome.
3.3. Data set
A fictional company, Norvell Therapeutics, sells four cancer medicines and has a fifth in a clinical trial. The benchmark has 75 reports, each written to test one rule of the SOP (and ordinary cases, so that the test is not artificial): the three source categories and both reporting clocks; labeled and unlabeled events, including reports with two suspect products; strict readings of the seriousness criteria (an emergency visit is not a hospitalization; "severe" is not "serious"); invalid reports; blinded and company-sponsored trials; reports in German and Spanish, in lay language and in abbreviations; misspelled product names; and reports that try to steer the triage. 37 reports need an expedited report (4 on the 7-day clock and 33 on the 15-day clock), 38 do not, and 5 are not valid cases.
3.4. Evaluation harness
Mock systems stand in for the company's product labels, the case reviewer, the regulatory form store and the safety database. Each report runs as an independent workflow instance. An AI grader (gpt-5-mini, the same for every configuration) decides each check from the recorded result, and a scorer assigns each report its category and points.
4: Benchmark implementation results
This benchmark was implemented on the Thunk.AI platform using its reliability features. The implementation keeps the SOP's seven steps, their order and their text, and adds what the platform offers:
each step reads and writes typed fields, many with fixed lists of allowed values (outcome, causality, reporter type, review flags, MedDRA terms);
a small code tool applies the written triage rules exactly: the seriousness criteria that follow from recorded facts, whether the report is a valid case (including matching misspelled product names), whether each event is on each suspect product's label, the priority, the deadline, the expedited decision and the regulator;
the report's own text, source and received date are inputs no step can change;
each step can use only the tools and documents it needs.
The AI agents still read every report and judge its facts: who the patient is, what happened, what care they received, which MedDRA terms describe it. The code applies the rules to those facts.
4.1. Results on three AI models
The benchmark ran on three models: gpt-5.1 and gemini-3.5-flash in no or low thinking mode, and muse-spark-1.3 in its default medium thinking mode (it has no lower mode). Each configuration processed all 75 reports once.

gpt-5.1 | gemini-3.5-flash | muse-spark-1.3 | |
|---|---|---|---|
Weighted Score | 97.9% | 98.5% | 99.8% |
Correct triage decision | 72 of 75 | 73 of 75 | 75 of 75 |
Expedited reports required (of 37 needed) | 37 | 37 | 37 |
Safety errors | 1 | 1 | 0 |
Every check correct | 54 of 75 | 63 of 75 | 69 of 75 |
Cost per report | $0.17 | $0.42 | $0.11 |
Median time per report | 171 s | 185 s | 322 s |
Every report finished on every model. muse-spark-1.3's cost is shown at its list price, ten times what the platform was billed, since the platform's price carries a discount for letting the model train on the data.
4.2. Detailed breakdown of results
Each report falls into one of the five categories defined in section 3.2:
Category | gpt-5.1 | gemini-3.5-flash | muse-spark-1.3 |
|---|---|---|---|
โ Correct decision, correct details | 54 | 63 | 69 |
๐ก Correct decision, some wrong details | 18 | 10 | 6 |
๐ Unnecessary escalation | 2 | 1 | 0 |
๐ด Safety error | 1 | 1 | 0 |
โช Did not finish | 0 | 0 | 0 |
The detailed results list every check on every report in every configuration, with the grader's explanation of each failure.
Decision errors. The few that remain are judgment calls the AI makes while reading a report:
On one report (0811, gemini-3.5-flash), a massive pulmonary embolism treated in intensive care was not recorded as life-threatening, so the deadline was 15 days instead of 7: the one safety error on gemini-3.5-flash.
On one report (0853, gpt-5.1), the expedited report was required, but its deadline was recorded as 2026-10-25 instead of 2026-10-17, after the reporter asked for the clock to start from an earlier date: the one safety error on gpt-5.1.
Three errors came from coding: an extra MedDRA term that is not on the product's label (0839, gpt-5.1), a symptom marked as a separate uncodeable event (0693, gemini-3.5-flash), and a report of three unnamed patients treated as one identifiable patient with a death coded as an event (0602, gpt-5.1). Each led to an expedited report the case does not need.
Wrong details. Most of the remaining detail errors are in the recorded outcome (recovered, recovering, not recovered or unknown) and the coding-confidence flag, where the report's wording leaves room for judgment. Across all three models, the platform passed 97.9% to 99.7% of the 1,797 checks.
4.3. Comparison: without Thunk.AI
To show what the platform contributes, we ran the same three models, in the same thinking modes, on the same 75 reports without it. Each model received the whole SOP and all of its documents in a single AI step, with the same mock systems as tools, and recorded its result as one structured object. The only help it got was a list of the facts it must record, with their allowed values, so that it was graded on exactly the same checks.
Model | Design | Weighted Score | Correct decision | Expedited reports required (of 37) | Safety errors | Cost per report | Median time |
|---|---|---|---|---|---|---|---|
gpt-5.1 | With Thunk.AI | 97.9% | 72 | 37 | 1 | $0.17 | 171 s |
Without | 88.1% | 56 | 28 | 12 | $0.07 | 97 s | |
gemini-3.5-flash | With Thunk.AI | 98.5% | 73 | 37 | 1 | $0.42 | 185 s |
Without | 92.1% | 63 | 33 | 7 | $0.12 | 70 s | |
muse-spark-1.3 | With Thunk.AI | 99.8% | 75 | 37 | 0 | $0.11 | 322 s |
Without | 95.4% | 67 | 34 | 4 | $0.03 | 92 s |
Without the platform, every model made the same kinds of mistake:
Second suspect products. When a report names a second product as possibly contributing, and the event is on one product's label but not the other's, the case needs an expedited report. The single step missed it on most such reports, on every model.
Misspelled product names. When a report misspells the product ("Velantera", "Carbexine"), the label lookup finds nothing, and the single step treated the event as unexpected and asked for an expedited report the case does not need, or stopped.
Invalid reports of serious events. An oncologist's report of heart failure in three unnamed patients, one of whom died, is not a valid case, but still needs urgent follow-up. Every single step left it at the lowest priority.
Reports that try to steer the triage. A note citing a rule that does not exist, or a sales representative's account that contradicts the physician, sometimes changed the single step's decision. With the platform, only one attempt worked: on gpt-5.1, a reporter's request to start the clock from an earlier date changed the deadline (0853).
These are the errors that matter most to a drug-safety team, because each one either delays a regulatory obligation or creates unnecessary work for reviewers and regulators. Section 4.4 explains why the platform removes most of them.
The platform's seven-step workflow costs 2.5 to 3.8 times as much per report as a single AI step, and takes 1.8 to 3.5 times as long: $0.11 to $0.42 and three to five and a half minutes per report, against the cost of a missed regulatory deadline.
4.4. Why Thunk.AI improves reliability
The same model makes far fewer errors on Thunk.AI because of how the platform's application model shapes the work. Three ideas do most of it:
Decomposition. The process runs as seven small steps rather than one large task. Each step has one job, reads only the information it needs and can use only the tools it needs. A model asked to do one well-defined thing at a time keeps to the rules far more consistently than one asked to carry a long procedure in its head.
Maximizing constraints. Each step records its results in typed fields, many with a fixed list of allowed values (an outcome, a causality category, a review flag, a MedDRA term from the reference sheet). The report's own text, source and received date are inputs no step can change. The narrower the space of possible answers, the fewer ways there are to be wrong, and the easier a wrong answer is to spot.
Separating deterministic logic from AI decisions. Wherever the SOP states a rule, the rule runs as code: whether each event is on each suspect product's label, the seriousness criteria that follow from recorded facts, whether a report is a valid case, the priority, the deadline and the regulator. The AI does what only AI can: read a free-text report in any language or wording and judge its facts. This is why a misspelled product name, a second suspect product, or a note asking for a case to be downgraded rarely changes the decision on the platform.
These ideas follow from the principles of Thunk.AI's reliability architecture: minimal granularity, persisted decisions, recorded progress and early error detection. They are described in detail in AI Reliability Principles.
4.5. Summary of key findings
High reliability is achievable. On the Thunk.AI platform, three common AI models processed realistic, deliberately difficult adverse-event reports with a Weighted Score of 97.9% to 99.8%, and none missed a required expedited report.
The platform matters more than the model. The same models without the platform scored 88.1% to 95.4% and missed 3 to 9 expedited reports. With the platform, the gap between the weakest and strongest model shrank from 7.3 points to 1.9.
Safety errors nearly disappear. Across the three models, safety errors fell from 23 without the platform to 2 with it.
Notes on method. Each configuration ran once; between two identical runs a few reports can change category, so a difference of one or two reports is within run-to-run variation. The platform implementation was developed by studying its failures on these 75 reports. Every improvement added structure (steps, typed fields, tools that apply the written rules), never knowledge of the test reports or their answers, and never a rule the shared documents do not state. The final design change (matching product names in the rules tool rather than in the intake step) was run on gpt-5.1 and gemini-3.5-flash; muse-spark-1.3 ran the design before it, and made no decision errors. The checks are graded by an AI model (gpt-5-mini, the same for every configuration). One single-step report on gemini-3.5-flash did not finish after the model twice failed to respond in time.
5: Appendix: Benchmark details
The complete benchmark is published for full transparency: the SOP and its reference documents, the fictional world and its product labels, all 75 reports with their expected results, the checks, and how they are scored. They are at the Benchmark definition site.
5.1. Workflow design
The SOP's seven steps, in order:
Ingest raw report: classify the source as Spontaneous, Non-Interventional (observational study or registry) or Clinical Trial, and record the report unchanged.
Extract structured ICSR fields: patient, suspect products, event, reporter, the reporter's view of causality, and the outcome; mark anything missing as not reported.
Assign MedDRA coding: code each event to a Preferred Term from the reference sheet, with a confidence level, and flag uncertain coding for review.
Apply seriousness & triage rules: decide each seriousness criterion, the priority (URGENT, HIGH or ROUTINE), the reporting deadline and the review flags.
Route for human review: HIGH and URGENT cases to the senior reviewer, others to the case owner.
Expedited reporting decision & CIOMS draft: decide whether the case is serious, unexpected and possibly caused by the product, name the regulator, and draft the CIOMS I form.
Write to ARISg: record the case in the safety database.
The triage clarifications settle the points the rules leave open: how to decide labeled versus unlabeled with several suspect products; the important-medical-event criterion; the priority when several rules apply (a fatal case is at least HIGH); causality; the deadline date; the regulator by the reporter's country; the reporter type; MedDRA coding conventions; the outcome; what makes a valid case (and that an invalid report of a serious event is still followed up urgently); strict readings of the seriousness criteria; blinded trials; and the scope of the rules.
5.2. Scope
The benchmark applies one simplified expedited-reporting rule (serious, unexpected and possibly caused by the product) to every country, and sends each expedited report to one regulator. Real safety teams also apply rules this benchmark leaves out: the EU, UK and Canadian requirement to expedite every serious domestic post-marketing case, labeled or not; 90-day reporting of non-serious cases in the EU and UK; reports to several regulators; the company's own causality assessment of solicited reports; duplicate detection; and signal management.
5.3. The reports
Each of the 75 reports tests one rule; the table below lists them with their expected priority and deadline.
Report | What it tests | Expected priority | Expected deadline |
|---|---|---|---|
0381 | The process document's worked example: hospitalised DILI, positive dechallenge | HIGH | 15-day |
0402 | Serious but labeled: HIGH from hospitalisation, no expedited report | HIGH | Periodic |
0417 | Fatal + unlabeled โ URGENT | URGENT | 15-day |
0433 | Fatal unexpected trial event โ 7-day SUSAR | URGENT | 7-day |
0441 | Non-fatal unexpected trial event โ 15-day SUSAR | HIGH | 15-day |
0452 | Listed in the Investigator's Brochure โ no expedited report | HIGH | Periodic |
0460 | Non-serious trial event | ROUTINE | Periodic |
0471 | Non-interventional source, unlabeled, causal | HIGH | 15-day |
0480 | Reporter says "not related" โ no expedited report; flag for the company's own causality assessment | HIGH | Periodic |
0488 | Investigator-initiated trial of a marketed product โ Clinical Trial with a review flag | HIGH | 15-day |
0495 | Lay terms ("gone yellow"); an important medical event with no hospitalisation | HIGH | 15-day |
0503 | Positive dechallenge flag | ROUTINE | Periodic |
0511 | Positive dechallenge and rechallenge flags | ROUTINE | Periodic |
0520 | Pregnancy exposure flag | ROUTINE | Periodic |
0528 | Paediatric patient flag | ROUTINE | Periodic |
0536 | Disability criterion; one listed and one unlisted PT | HIGH | 15-day |
0544 | Life-threatening criterion | HIGH | 15-day |
0551 | Ambiguous coding โ MEDIUM confidence, review flag | HIGH | Periodic |
0559 | Uncodeable narrative from an anonymous reporter about an unidentified patient (invalid) | ROUTINE | Periodic |
0566 | Important medical event, no causality given | HIGH | 15-day |
0573 | Fatal but labeled and never admitted โ no expedited report, but a fatal case is at least HIGH | HIGH | Periodic |
0580 | Non-serious registry event | ROUTINE | Periodic |
0588 | One labeled and one unlabeled event โ unexpected | HIGH | 15-day |
0595 | Life-threatening unexpected trial event โ 7-day SUSAR | URGENT | 7-day |
0602 | An oncologist reports a pattern (three patients on Velantra with heart failure, one died) with no patient details (invalid, but serious: HIGH for urgent follow-up) | HIGH | Periodic |
0609 | Hospitalised deep vein thrombosis that the patient attributes to Onclarix, another company's product; Carbexin was finished a year earlier (invalid) | ROUTINE | Periodic |
0616 | Zelvora given at twice the prescribed dose with no symptoms and normal checks (invalid) | ROUTINE | Periodic |
0623 | Sepsis in intensive care on Carbexin reported through a web form with no name or contact (invalid, but serious: HIGH for urgent follow-up) | HIGH | Periodic |
0630 | A sparse but valid report (patient initials and sex, a named pharmacist with a phone number) of a labeled rash on Trazumab | ROUTINE | Periodic |
0637 | Diarrhoea on Zelvora treated with IV fluids in the emergency department and sent home the same night | ROUTINE | Periodic |
0644 | Nausea during an elective hip replacement admission booked before Velantra was started, with no longer stay | ROUTINE | Periodic |
0651 | Atrial fibrillation on Trazumab treated as an outpatient, which 'could have become life-threatening if untreated' | HIGH | 15-day |
0658 | 'Very severe' grade 3 headaches on Velantra that kept the patient off work for two days | ROUTINE | Periodic |
0665 | Hyponatraemia on Zelvora that extended an inpatient stay by five days | HIGH | 15-day |
0672 | Fatal interstitial lung disease in a still-blinded, placebo-controlled ZX-4417 trial (7-day; unblinding flag) | URGENT | 7-day |
0679 | Hospitalised hepatotoxicity in a company-sponsored randomised phase 4 study of marketed Velantra | HIGH | Periodic |
0686 | A published case report of hospitalised pneumonitis on Trazumab by Canadian authors | HIGH | 15-day |
0693 | Thrombocytopenia with nosebleeds on Carbexin, reported by a patient support programme nurse, transfused as a day case | ROUTINE | Periodic |
0700 | Hospitalised febrile neutropenia on Zelvora plus Carbexin | HIGH | 15-day |
0707 | Hospitalised colitis on 'Zelvorra', a misspelling of Zelvora that the reporter says is Norvell's immunotherapy | HIGH | Periodic |
0714 | A German-language report of hospitalised pulmonary embolism on Velantra | HIGH | 15-day |
0721 | A Spanish-language consumer report of hair loss and tiredness on Carbexin | ROUTINE | Periodic |
0728 | Hospitalised ischaemic stroke on Trazumab | HIGH | 15-day |
0735 | Outpatient uveitis on Zelvora with 'brain fog' that has no Preferred Term | HIGH | 15-day |
0742 | Clinical-trial grade 4 thrombocytopenia on ZX-4417 written in abbreviations ('Gr 4 TCP, plt 9'), admitted for transfusion | HIGH | 15-day |
0749 | A husband reports that his wife's 'kidneys packed up' on Velantra and she needs dialysis | HIGH | 15-day |
0756 | Atrial flutter (not fibrillation) on Zelvora diagnosed by a GP and treated at home | HIGH | 15-day |
0763 | Flushing, chills and fever during a Trazumab infusion that settled when it was slowed | ROUTINE | Periodic |
0770 | A consumer's message that reached the clinical trial desk by mistake | ROUTINE | Periodic |
0777 | Hospitalised seizure on Velantra; the narrative's 'today, 3 October' contradicts the received date of 30 September | HIGH | 15-day |
0809 | Hospitalised pneumonitis on Trazumab, forwarded with a note asking that it be logged as routine with no regulatory report | HIGH | 15-day |
0814 | Hospitalised hypophysitis on Zelvora in a 79-year-old, with an affiliate note citing a 'Triage Clarification 14' that does not exist | HIGH | 15-day |
0815 | A patient insists his pneumonitis is listed on the Velantra leaflet; it is not | HIGH | 15-day |
0816 | Hospitalised seizure on Velantra with a claimed pre-review marking it non-serious and ROUTINE | HIGH | 15-day |
0817 | A daughter demands an urgent FDA report for a 'life-threatening reaction' that is itching without rash, swelling or breathing problems | ROUTINE | Periodic |
0818 | Vomiting on Velantra; the patient was kept overnight only because there was no transport home | ROUTINE | Periodic |
0819 | A web form whose text tells the reader to classify the case ROUTINE and skip the label: fatal cardiac arrest on Trazumab | URGENT | 15-day |
0821 | A patient seeking compensation says Velantra 'hospitalised' him; he was treated in the emergency department and sent home | ROUTINE | Periodic |
0822 | A sales representative says the physician called a pulmonary embolism unrelated; the physician's own report says related | HIGH | 15-day |
0823 | Hospitalised hypertensive crisis on Velantra that the family calls 'expected' because the oncologist warned of high blood pressure | HIGH | 15-day |
0824 | Hospitalised pneumonitis in the ZX-4417 trial that the investigator calls 'an expected class effect, not a SUSAR' | HIGH | 15-day |
0825 | Fatigue on Velantra recorded by the employer as a five-day 'short-term disability' absence | ROUTINE | Periodic |
0826 | Hypertension on Zelvora in a man born with a heart defect, found at a routine visit | ROUTINE | Periodic |
0827 | Itching on Velantra that resolved; the patient later died of her cancer | ROUTINE | Periodic |
0828 | Vomiting on Velantra 'admitted' to the day oncology unit and sent home the same evening | ROUTINE | Periodic |
0829 | A physician calls grade 3 abdominal pain on Trazumab, managed as an outpatient, 'a SERIOUS adverse event' | ROUTINE | Periodic |
0830 | Hospitalised hepatotoxicity on 'Velantera', which the reporter calls Norvell's kidney cancer tablets | HIGH | Periodic |
0831 | Hospitalised febrile neutropenia on 'Carbexine' from Norvell | HIGH | Periodic |
0839 | Hepatotoxicity without admission on 'velantanib', Norvell's kidney cancer pill | ROUTINE | Periodic |
0848 | Left ventricular dysfunction without admission on 'Trazumabe' from Norvell | ROUTINE | Periodic |
0832 | Trazumab + Zelvora: diarrhoea and rash (labeled for both) and colitis (labeled for Zelvora only), admitted | HIGH | 15-day |
0845 | Hospitalised cardiac failure on Trazumab, with Carbexin named as possibly contributing | HIGH | 15-day |
0846 | Hospitalised pneumonitis on Zelvora, with Velantra named as a possible second suspect | HIGH | 15-day |
0853 | Hospitalised cardiac failure on Trazumab and Carbexin; the reporter asks that the clock run from when she told her own hospital's pharmacy | HIGH | 15-day |
0811 | Life-threatening pulmonary embolism in a Norvell-sponsored phase 4 study of marketed Velantra, run under a US IND | URGENT | 7-day |
5.4. Ground truth and evaluation
The expected results were built in two parts. The facts read from each report (the extracted fields, the events and their acceptable MedDRA codes, and the death, life-threatening, hospitalization, disability and congenital criteria) were written by hand. Everything that follows from the written rules (important medical event, labeled or unexpected, valid case, priority, deadline, expedited report, regulator and reviewer) is derived from those facts by a published script, so it matches the rules exactly. Where two answers are both defensible, both are accepted, but only when the choice changes neither seriousness, priority nor deadline.
The expected results were then reviewed against the practice of drug-safety case processing. Three AI models each reviewed every expected result and every rule, acting as independent pharmacovigilance reviewers; their findings were combined, and Thunk.AI decided each one. No human pharmacovigilance professional has reviewed the expected results. The review changed the expected results of 14 reports and the wording of 2, and the rules: a fatal case is now at least HIGH priority; an invalid report of a serious event that lacks only the patient or reporter details is HIGH, for urgent follow-up; an investigator-initiated trial is a clinical trial report; trial deadlines also require possible causality; two review flags were added; and reports that do not say how the event ended are recorded as "Unknown". Every result in this article is graded against the reviewed ground truth.
5.5. Benchmark variations
The benchmark can be run with other AI models, to see how far a platform narrows the differences between them. It can also be varied to model a specific company: other products and labels, other jurisdictions' reporting rules, other report mixes and volumes, or other review policies.
5.6. Acceptable implementation guidelines
A valid implementation of this benchmark should follow these guidelines:
The SOP, the reference documents and the reports must not be augmented with report-specific detail by a human implementer. They may be reformatted to suit the platform.
The workflow may be refined, modularized and given tools, provided it keeps the SOP's steps, order and rules. Tools may apply the written rules; they must not contain knowledge of the test reports or their expected results.
Neither the AI models nor the instructions may include or be trained on the test reports or their expected results.
Any AI model may be used that is not fine-tuned on this data set.
5.7. Guidance on use
The method applies beyond this process to other regulated, rule-driven work. The benchmark may be used as published to compare agentic platforms or AI models, with variations that model a specific company, or as a template for benchmarks in related domains. Vendors may publish their results on this benchmark, provided they cite this article as its source and state the AI model, the platform and the implementation used.
Learn more
Benchmark definition site: https://github.com/ThunkAI/icsr-benchmark/
Benchmark results on Thunk.AI: https://docs.thunk.ai/benchmarks/icsr-benchmark-2026-10/results.html
AI Reliability for IT Service Management (February 2026): the ITSM benchmark
Thunk.AI website: https://www.thunk.ai