Chile manages 400,000+ surgical waiting list records monthly. Cancer patients often remain invisible within this volume — despite facing 7.5x higher mortality than the rest. We built a sovereign in-house AI that screens oncological risk across all 29 national Health Services, flagging a high-risk group in hours instead of weeks. Sensitivity: 95%. Manual review burden cut 99.76%. No vendor, no black box: AI flags risk, humans decide. Every patient, every region, the same standard.
Innovation Summary
Innovation Overview
The Problem
Chile's public health system processes approximately 400,000 active surgical waiting list records each month across 29 Health Services. Hidden within this volume, there are patients with suspected cancer who urgently need prioritization. The challenge is not scale alone: records are written in inconsistent free text with local abbreviations, misspellings, and non-standardized terminology that varies across every service. Manual review is slow, inconsistent, and geography-dependent — a patient in a remote region is statistically far less likely to be flagged than one near a major hospital. The result is invisible inequity embedded in the administrative process itself — and a measurable human cost: cancer-suspected patients, representing just 3.8% of the waiting list, carry 7.5x higher mortality than the rest.
The Innovation
We developed a fail-safe hybrid AI system that classifies oncological risk across three independent layers: (1) clinical heuristics encoding expert rules; (2) semantic embeddings capturing natural language meaning despite terminological variation; and (3) an XGBoost statistical classifier trained on nationally validated data. When layers conflict, the case is automatically escalated to mandatory human review — AI assists, but it never decides. The system is built entirely in-house on sovereign infrastructure, with no external vendor dependency and full decision auditability.
Why It Simplifies Government
This innovation converts an unmanageable administrative bottleneck — reviewing 400,000 heterogeneous free-text records monthly — into a standardized, automated, and continuously measured national process. It directly eliminates one of the most burdensome tasks in Chile's public health administration, freeing clinical staff to exercise judgment rather than perform screening.
Beneficiaries
Cancer-suspected patients receive timely identification — and the data shows it matters: detected patients access surgery 4.8x faster than undetected ones (206 vs 986 days) and are resolved 4.6x sooner overall (157 vs 720 days; p<0.001). Review burden was dramatically reduced. Health network managers gain national traceability of the oncological caseload. For the first time, a patient in Chile's southernmost region receives the same oncological screening standard as one in Santiago.
Replicability
The multilayer fail-safe architecture is transferable to any health system managing large volumes of heterogeneous administrative records — particularly relevant for middle-income countries across Latin America, Sub-Saharan Africa, and Southeast Asia facing challenges of scale, data heterogeneity, and limited specialist capacity.
Future
The system is transitioning from retrospective analysis to a real-time Quality Layer integrated at hospital level — an invisible auditor preventing patients from falling through administrative cracks at the point of record entry.
Innovation Description
What Makes Your Project Innovative?
Three things set this project apart.
- Technically: three independent layers — clinical heuristics, semantic embeddings, XGBoost — create a fail-safe architecture. If the statistical model degrades, the system stays stable. Uncertainty is managed, not hidden.
- Ethically: fully sovereign, built in-house on Chile's own data, no external vendor. Every decision is explainable (SHAP/LIME) and subject to mandatory human review before any administrative action. Human-in-the-Loop is not a feature here — it is the foundation.
- Most unexpectedly: a virtuous data quality cycle emerged. As clinical teams received AI feedback on their records, they began recording more carefully — precision rose from 71.5% to 81.1% in six months, entirely unplanned. Well-designed public AI can improve the ecosystem it depends on, not just its outputs.
This generalizes to any government domain where data heterogeneity creates administrative bottlenecks.
What is the current status of your innovation?
Implementation and Evaluation. The system has been operating for over 9 months across all 29 Health Services in Chile, processing the full national surgical waiting list on a weekly cycle with continuous performance monitoring. Sensitivity, specificity, PPV and NPV are tracked per Health Service, enabling real-time detection of quality degradation. Institutionalization is underway to transition from retrospective analysis to a prospective real-time Quality Layer integrated at hospital level.
Innovation Development
Collaborations & Partnerships
Developed by the Advanced Analytics Unit of Chile's Ministry of Health (MINSAL), with active collaboration from multidisciplinary teams across all 29 Health Services nationwide. These teams serve as clinical validators of the model's ground truth — their distributed expertise was essential not only for model accuracy but for building the institutional trust that enabled national adoption. Future phases will incorporate clinical teams as co-designers of validation criteria, not just validators.
Users, Stakeholders & Beneficiaries
- Patients: Timely identification of cancer-suspected individuals reduced critical delays in a system, where waiting times span months.
- Clinical staff: Dramatically reduced screening burden, enabling focus on clinical judgment over record review.
- Medical specialists: Better-informed prioritization of surgical cases.
- Health network managers: National visibility, traceability, and a standardized oncological risk signal — comparable across all regions for the first time.
Innovation Reflections
Results, Outcomes & Impacts
Across 732,837 cases tracked over 9 months, the system flags a high-risk group — just 3.8% of the waiting list — carrying 7.5x higher mortality (33.1 vs 4.4 per 1,000; p<0.001). Flagged patients access surgery 4.8x faster (206 vs 986 days), resolve 4.6x sooner (157 vs 720 days; p<0.001), and reach surgery more often (53% vs 36.3%). Sensitivity: 95%; NPV: 99.3%. Manual review burden cut 99.76%. A virtuous cycle emerged: as local Health Services received AI feedback, they improved data quality at the source — driving PPV from 71.5% to 81.1% and strengthening local governance of oncological recordkeeping across all 29 services. Every patient, every region, same standard.
Challenges and Failures
The hardest challenge was not technical — it was human. Initial teams across several Health Services applied flawed validation logic, comparing retrospective AI outputs against prospective clinical records, producing misleading metrics and unfounded distrust. A structural failure persists: 23.2% of flagged patients remain unvalidated. The impact report shows these cases carry worse survival outcomes — detection alone is insufficient if the system cannot process what it finds.
Our response: a mandatory National Validation Protocol, regional multidisciplinary teams to realign ground truth criteria, and weekly per-service dashboards to make gaps visible. Governance had to be rebuilt alongside the model, not after it.
Conditions for Success
- Infrastructure: Sovereign in-house infrastructure gave full control over health records and enabled rapid iteration without vendor constraints.
- Policy: A national mandate ensured adoption across all 29 Health Services. Local opt-in would have produced a fragmented, unreliable system.
- Leadership: Ministerial support converted a technical tool into an institutional priority — critical for sustaining adoption through early resistance.
- Resources: A small interdisciplinary team combining clinical knowledge, NLP, and public health expertise proved more effective than scale alone.
- Motivation: Every design decision was driven by one principle: no patient should be invisible to the system meant to find them.
Lessons Learned
In public AI, governance matters as much as code. The initiative brings key lessons. Define ground truth collaboratively before building the model. Involve frontline staff as co-designers, not just validators — their resistance often reflects legitimate methodological disagreement. Build explainability in from day one: it converts skeptics into advocates. One of the most significant findings — that AI feedback improves data quality at the source — was unplanned and carries broad implications for AI governance globally. Measurement is a feature, not an afterthought. For any government seeking to replicate: build governance first, then the model, then the metrics. The architecture matters less than the process surrounding it.
Anything Else?
Before this system existed, only 3.8% of patients carrying 7.5x higher mortality was invisible — indistinguishable from 400,000 other records in a system without the capacity to review timely each case. They waited. Some did not make it. This is not primarily a technical achievement. It is a governance argument: that a small public sector team can build ethical, sovereign AI — no black box, no vendor, every decision explainable and subject to mandatory human review. Human-in-the-Loop is not a feature here; it is the foundation. AI flags risk. Humans decide. That equity is an engineering requirement, not only a value. The architecture is replicable. Any government willing to treat administrative data as a clinical asset can build this.
Supporting Videos
Status:
- Implementation - making the innovation happen
- Evaluation - understanding whether the innovative initiative has delivered what was needed
- Diffusing Lessons - using what was learnt to inform other projects and understanding how the innovation can be applied in other ways
Date Published:
25 September 2026

