The Hybrid Intelligence Model for Risk Assessment and Post-Market Surveillance

How Smarteeva Orchestra Divides Work Between Formulas and Language Models
Automating a customer's complaint handling risk assessment taught us something we did not expect to be the main lesson.
The strongest reporting systems do not pick a side between deterministic computation and large language models. They combine the two on purpose, and the value comes from where the boundary sits.
Deterministic methods give precision, consistency and traceability for quantitative information. Language models give the contextual reasoning needed to classify records, recognise patterns and write something a person would want to read. Getting the division right is what makes automated risk assessments, Periodic Safety Update Reports and other post-market surveillance reports possible without giving up the rigour a validated system requires.
Assigning each job to the mechanism suited to it
The risk assessment was generated using Smarteeva Orchestra AI Agent Orchestrations, which coordinates multiple models and supporting processes.
The implementation did not treat the language model as a universal solution. Each type of work went to the mechanism best able to perform it.
Quantitative facts came from deterministic queries and formulas. That covers complaint counts, affected product populations, sales volumes, installed base figures and calculated complaint rates. Classification, quality interpretation, trend characterisation and report writing went to the models.
That distinction turned out to be the whole design.
Why numbers must be deterministic
Numbers in a regulated report have to be reproducible.
Suppose a complaint rate is calculated from 47 qualifying complaints across an installed base of 22,000 devices. The same inputs and the same formula must always produce the same result. It cannot depend on how a prompt was worded, which model answered, or how a probabilistic system interpreted the question that day.
A reviewer also has to be able to trace the result backwards, to the source records, the inclusion criteria, the calculation logic and the reporting period.
So deterministic computation is the right mechanism for counting records, applying date boundaries, filtering product families, aggregating sales or installation data, and calculating rates or period-over-period changes. These operations can be specified, tested, validated and run repeatedly with predictable results. In a validated environment that predictability functions as a control, which is a higher bar than convenience.
Why classification is a different problem
Complaint classification does not behave like arithmetic.
Complaint records arrive with unstructured narratives, inconsistent terminology, incomplete descriptions, and context scattered across several fields. Two complaints describing the same issue may share no vocabulary at all. Two complaints using similar words may describe events with different causes, severities and outcomes.
Exact matching fails here. Rigid rule sets fail here. The work requires contextual interpretation.
This is where the model-driven part of the orchestration earned its place. In this implementation, the models consistently outperformed human agents at classifying each complaint.
The reason matters more than the claim. The orchestration evaluated every eligible record against a common taxonomy rather than a sample, applied the same analytical framework from the first record to the last, and did not tire across thousands of narratives. Manual review under time pressure cannot offer any of those three things, which is not a comment on the people doing it.
The controls that make it defensible
The advantage does not remove the need for controls.
Model-assisted classification stays traceable to the source complaint, the applicable taxonomy, the model and prompt configuration, and the rationale behind the assigned category. Confidence thresholds and exception handling then direct human review toward ambiguous and high-risk cases.
That produces a better use of expert judgment. Specialists oversee the analytical framework and investigate meaningful exceptions, instead of spending most of their week performing repetitive classification.
The same division applies to writing the report
A language model is well suited to synthesising findings, describing trends, connecting related evidence and producing a coherent risk narrative.
It can assess whether an observed increase looks meaningful, summarise the dominant complaint categories, and draft conclusions in language appropriate to a formal report.
What it must not do is invent, estimate or independently recalculate the underlying facts. The model writes from an approved body of evidence that deterministic processes assemble.
The chain of responsibility
- Deterministic processes establish the quantitative record
- Language models classify and interpret the qualitative information
- Deterministic controls verify totals, rates, thresholds and internal consistency
- Language models convert verified evidence into a structured analytical narrative
- Qualified reviewers assess exceptions, conclusions and final approval
The outcome is faster document production, and underneath that, a more disciplined form of automation where each component does work aligned with what it is good at.
Why this matters most in post-market surveillance
Post-market surveillance reporting combines structured measures with complex qualitative interpretation more than almost any other regulated document.
A PSUR may require complaint and incident counts, exposure or sales information, reporting rates, trend comparisons, benefit-risk conclusions, summaries of corrective actions and evaluation of emerging signals.
Build all of that with deterministic rules alone, and you get a brittle system that breaks on nuanced language and shifting real-world evidence. Build it entirely with a language model, and you introduce risk around numerical accuracy, reproducibility and traceability, which are the three properties a regulator examines first.
A hybrid system avoids both failure modes. Calculations stay controlled computations. Language understanding stays a reasoning task.
It can reconcile report totals against source systems, preserve the lineage of every metric, and stop unverified numerical statements from entering the narrative. At the same time, it can analyse complete complaint populations, identify semantically related events, and produce reports more consistent than fragmented manual workflows allow.
The boundary is the design decision
Successful AI adoption in a regulated environment depends less on how much work you can hand a language model, and more on where you draw the line between probabilistic and deterministic processing.
Models should not perform arithmetic a formula executes exactly. Deterministic systems should not be forced to interpret narratives that need contextual reasoning.
Formulas and queries establish what happened. Models help explain what it means. Validation controls keep the finished report reliable, traceable and fit for regulated use.
If you are scoping AI for your own quality reporting, start by listing every output in the report and marking each one as a calculation or an interpretation. That list is your architecture.
FAQs
- Can AI generate medical device regulatory reports? Parts of them. A hybrid architecture works best: deterministic queries and formulas produce every quantitative fact, language models classify records and draft the narrative, and qualified reviewers approve the conclusions.
- Should a language model calculate complaint rates? No. Rates must be reproducible from the same inputs every time and traceable to source records, inclusion criteria and reporting period. That requires deterministic computation rather than probabilistic interpretation.
- Why use a language model for complaint classification at all? Complaint narratives are unstructured and inconsistent. Two records describing the same issue may share no vocabulary, and similar wording can describe different causes and outcomes. Classification needs contextual interpretation, which exact matching and rule sets cannot provide.
- How is model-assisted classification kept auditable? Each classification stays traceable to the source complaint, the applicable taxonomy, the model and prompt configuration, and the rationale for the assigned category. Confidence thresholds route ambiguous and high-risk records to human review.
- Does this remove people from regulatory reporting? No. Qualified reviewers assess exceptions, conclusions and final approval. What changes is that specialists oversee the analytical framework instead of performing repetitive classification.






