The programme counted 17,415 visits to show that doctors saved time. It asked six patients what it felt like. That gap is the point of this page.
This page reads the public documents on Ontario's AI scribe deployment. Patients were promised they would be told, asked for consent, and able to check or correct what was written. The question is what the public record actually shows.
Positioning
This is an empirical case study built on public records rather than user testing. It examines how one institution decided which evidence to produce, and whose interests were made measurable. The paper's architecture is laid out in S7.
An AI scribe listens to a doctor–patient visit, creates a transcript, and drafts a clinical note for the electronic medical record. In Ontario, the Ministry of Health told CBC in May 2026 that roughly 5,000 physicians were using one — a ministry figure, not an audited programme statistic. The products come from Supply Ontario's Vendor of Record arrangement (Tender-20123), which lists 30 qualified vendors after a second intake; 20 were approved when it was awarded in April 2025.
The programme's evaluation report measured clinician benefits in detail. Between March 18 and July 5, 2024, 152 primary care providers logged 17,415 encounters; in the lab, documentation time fell 69.5% across nine simulated encounters (Table 1), and in routine practice providers self-reported about three hours less administrative work per week (Figure 2). The number most often quoted in public — 70% to 90% less time on paperwork — is not in that report. It comes from OntarioMD's news release. For patient experience, the same evaluation relied on semi-structured interviews with six patients (n = 6).
Three instruments set the expectations: CPSO's Advice to the Profession: AI Scribes in Clinical Practice (June 2024), Ontario's health privacy law (PHIPA), and the Information and Privacy Commissioner's AI Scribes: Key Considerations for the Health Sector (January 28, 2026). CPSO advice is companion guidance to a policy — Medical Records Documentation — rather than a free-standing rule, which is itself part of what makes this mould hard to test:
Informed Communication
Physicians are expected to tell patients how the AI scribe will be used for documentation, in language the patient can act on.
Explicit Consent
Consent must be obtained before recording an encounter. The IPC treats express consent recorded in the chart as best practice, and says a patient who declines must receive the same standard of care.
Substantive Review
Physicians must review all generated notes for accuracy and completeness before saving them to the chart.
Privacy Protection
Systems must safeguard personal health information across storage, transfer, and any cross-border processing.
Bias & Suitability
Procurement required solutions to handle diverse accents, multiple speakers and clinical acronyms; professional guidance asks physicians to stay alert to bias and to whether a tool suits their patient population.
Continuing Responsibility
AI assists but never replaces professional clinical judgement; the physician remains solely accountable for the record.
The same public facts can be read as a benefit, a use of persons, or a habit of character. The three theories do not need a new dataset. They need a different unit of moral concern: the sum, the person, the character being formed.
Utilitarianism
Which action produces the most good for the most people?
The programme already ran half of that sum. Lab documentation time fell 69.5%; in the field, providers self-reported about three hours less a week over 17,415 visits. That is real good — for clinicians, and possibly for patients who get the time back as attention. It is not a completed calculation. Patient-side utility rests on six interviews. Treating "69.5% versus n=6" as a single ratio is rhetoric, not arithmetic: one number is a lab time cut, the other is a field sketch. A utilitarian still has to count later harms that reach anyone — fabricated detail, a wrong drug, a missed mental-health line. The honest claim is narrower: one stakeholder was instrumented; the other was not. The theory does not forbid the scribe. It forbids declaring the case closed while a column is empty.
Kantian ethics
Are persons treated as ends, or only as means?
Kant's formula of humanity: a person may be used as a means — a clinical encounter always is — but never merely as a means. They must also be able to know, refuse, and have a say in what is made of them. The evaluation turned those encounters into the evidence that doctors saved time. Six of the people in the room were asked what it felt like. The rest of the recorded patients functioned as the generation mechanism for an efficiency statistic. That is the charge: not that efficiency was measured, but that the persons whose speech was captured were not treated as ends in the same act that used them as data. "The doctor will review" does not restore that status. It is a label held by a third party. A Kantian programme could still cut documentation time. It would have to make refusal costless, the note inspectable, and correction possible — otherwise the patient remains a means of producing relief for someone else.
Virtue ethics
What kind of physician is this practice forming?
The virtue at stake is not speed. It is remaining the author of the record: conscientiousness about what one has actually read, honesty about what one has not. CPSO requires substantive review before the note enters the chart. Under back-to-back appointments that duty can shrink to a formal signature — and none of the approved systems required an attestation step (Auditor General, p. 24). This is not evidence that physicians are rubber-stamping. It is evidence that the workflow makes the signing physician easier to become than the reading one. The hours the 69.5% freed are also the hours a certain kind of physician would spend staying the author. The same question sits one level up. An institution that instruments clinician hours and leaves patient safeguards as principle is becoming a particular kind of steward.
Comparing the public evaluation against what remains unverified reveals a consistent pattern:
| Publicly visible | Less clear in public materials | Finding it supports |
|---|---|---|
| 152 providers, 17,415 encounters, 69.5% less documentation time in the lab, about three hours a week self-reported, satisfaction surveys | Error and omission rates across diverse accents, languages, multi-speaker visits, or sensitive conditions | Efficiency and workload relief are far more visible than clinical safety or equity data. |
| Official statements that the qualified vendors meet provincial privacy and security standards | Specific test criteria, pass thresholds, exception policies, and ongoing audit procedures | Procurement qualification claims cannot be independently verified by patients or external researchers. |
| Professional guidance requiring informed patient consent and mandatory doctor review | How consent is recorded in practice, refusal rates, and whether rushed workflows permit thorough note checks | A formal rule that "doctors review" does not guarantee that thorough error-checking happens during busy clinic hours. |
| Accuracy testing of 20 systems, run by procurement evaluators in 2024 and made public only by the Auditor General's special report of May 12, 2026 | Routine monitoring by vendors or the programme to catch fabricated detail before a note enters the chart; no approved system required a physician attestation step | The accuracy evidence existed inside the procurement two years before the public saw it, and reached them through an external audit rather than routine reporting. |
Paris, Moon and Guo (FAccT '26, pp. 5348–5370, DOI: 10.1145/3805689.3812373) examine how technical verification rests on unexamined assumptions. Drawing on the misabstraction framework of de Troya et al., they identify four fallacies that cause misabstractions in AI verifiability work (Fig. 1, p. 5356). Two of them carry this analysis: §4.1 and §4.3. The other two are readable here but are not what the argument rests on.
§4.1 Non-contextual claim
The forum's expectations — that a note is accurate, that a person is safe — have to be turned into a claim that can be checked. Here “qualified on the vendor list” was defined, weighted and scored by the procuring body alone, as though there were no argument about what quality or privacy mean (p. 5356).
§4.2 Sequence of computations
Assuming that because transcription and summarisation verify separately, the combined documentation system is verified — setting aside the human factors and the interdependencies between components (p. 5356).
§4.3 Trustless
Patients are positioned as able to check the claim while lacking the access, the resources and the standing to do so. The accuracy results stayed inside the procurement for two years, so a patient had to trust Supply Ontario before there was anything to examine (p. 5356, p. 5359).
§4.4 Narrow threat
Treating the distrusted actor as able to tamper only with the AI system, while it holds structural power outside it (p. 5356). Readable here — the procurer set the weights and required no attestation step — but the pair above gives the sharper account.
Whose definition?
What “safe” and “high quality” might mean to a patient
- my chart does not record the wrong medication
- I do not get the wrong treatment because of an error in an AI note
- my doctor really has the time, the duty and the institutional requirement to review the note
- I can understand what is being recorded, and I can say no without losing care
- if the record is wrong, I can find out, object, and get it corrected
- if a problem is found in the system, I am told and something is done about it
What the procurement compressed this into
- the system passed a number of accuracy tests
- the vendor holds certain privacy or security certifications
- the vendor reached the pass mark in the procurement scoring
- the vendor is on the approved vendor list
So the problem is not that the procurement body has no definition. It is the opposite. The problem is that it holds the power to define on its own. A normative question that patients, doctors, patient organisations, regulators and the public should decide together becomes a set of technical and administrative measures that one body can manage internally.
The question this page is working toward
When a public procurement body verifies an AI system in the name of patients, but patients cannot take part in setting the criteria, cannot obtain or interpret the evidence, and cannot challenge the conclusion in time, is that verification a legitimate form of AI accountability from the patient's point of view?
The dimensions to work through:
| Dimension | Question | Material in the Ontario case |
|---|---|---|
| Forum | Who is actually affected, and who is appointed to represent them? | patients, doctors, patient organisations, Supply Ontario |
| Distrusted actor | From the patient's side, who has to answer for this? | vendors, Supply Ontario, or the two together as a procurement chain |
| Claim | What is being claimed? | “qualified vendor list”; the systems meet quality, privacy and security requirements |
| Operationalisation | Who decides what the claim means and how it is measured? | procurement criteria, clinical evaluation, the design of the accuracy tests |
| Evidence access | Who can see the raw tests, the results, the limits and the failure modes? | procurement staff, vendors, the audit office, patients |
| Interpretability | Who has the resources and the ability to understand the evidence? | clinical experts, procurement staff, patients and their representatives |
| Challenge power | Who can question it, ask for an explanation, change the criteria, or trigger a remedy? | whether patients have any real channel; whether audit can only act afterwards |
| Timing | When does the evidence become public? | accuracy problems found during procurement in 2024, made public by external audit in 2026 |
| Fallacy mapping | What is the main failure? | non-contextual claim and trustless, rather than narrow threat |
Approach
Document-based analysis of publicly available materials: the WIHV/CDHE clinical evaluation report (July 31, 2024), the Supply Ontario Tender-20123 record, CPSO advice and the Medical Records Documentation policy, the IPC's AI scribe guidance (January 28, 2026), OntarioMD programme material, and the Auditor General's Use of Artificial Intelligence in the Ontario Government (May 12, 2026). Full list in the references.
Evidence Coding
Each patient-facing commitment in the governing framework is evaluated against the public record and coded into one of three categories: evidenced (supported by public data), asserted (stated as policy without verification data), or absent (no public documentation).
What This Project Is Not
This project does not conduct clinical trials or evaluate proprietary model code. It does not argue that AI scribes are ineffective or that doctors should not use them. The focus is strictly on the shape of the public evidence base.
The Bounded Finding
The programme built a measurable case for efficiency, while patient protections remained principles.
The public record shows clear, quantified evidence that AI scribes reduce clinician administrative burden, gathered across a deployment of 17,415 encounters. For patient-facing commitments — transparency, consent, review, and correction — the public record largely contains assertions rather than operational evidence. This does not mean patients were harmed; it shows that the programme was designed to measure provider success, not patient safeguards.
For the pitch, this page moves from a conceptual critique to an empirical case study. Paris, Moon and Guo examine a body of technical papers; this project examines one deployment. Their own discussion asks for exactly that: more empirical research and case studies of the practices and norms around verification (§5.1.2, p. 5361). The paper keeps their section order, but each section now carries evidence from the Ontario record instead of from prior literature.
Research questions
RQ: Across Canada's AI scribe programs, how do governance and procurement arrangements translate patient safety and clinical trustworthiness into checkable criteria, and how do different procurement models distribute the responsibility those criteria leave unaddressed?
The case widens from Ontario to four programs. The question in S5, whether verification done in patients' name is legitimate accountability from their side, remains the evaluative lens for the answers.
- Sub-RQ1 (documents): How do the governance and procurement instruments of Ontario's AI scribe deployment and Canada Health Infoway's national program operationalise patient safety and clinical trustworthiness as checkable criteria, and which patient-facing expectations do those criteria leave out?
- Sub-RQ2 (documents): How do centralised vendor vetting (Ontario), a single-vendor mandate (Nova Scotia), and direct clinician–vendor contracting (British Columbia) assign residual responsibility, and what mechanisms in the public record allow that responsibility to be discharged or checked?
- Sub-RQ3 (public online discussion): How do Canadian clinicians describe review, consent and error correction in practice in public online discussion, and where do those accounts diverge from the assumptions in the documents?
Sub-RQ3 depends on a McGill REB determination and on a permitted route to Reddit data; its limits are listed in the REB prep document below.
| Paris et al. (conceptual critique) | This case study | What the Ontario record supplies |
|---|---|---|
| §1 Introduction. Firms have incentives to evade scrutiny (Dieselgate); technical work offers computational verification as the fix. | 1 Introduction. One deployment in which clinician relief and patient safeguards were documented unevenly. | 17,415 encounters and a 69.5% lab result, against six patient interviews; the 70–90% figure that is not in the evaluation report. |
| §2 Background. Accountability without trust (Bovens, §2.1); access as a limited solution; the technical landscape (ZKP, TEE, watermarking). | 2 Case context and governance baseline. Who set the expectations, and what vendors were qualified against. | Tender-20123; CPSO advice and the Medical Records Documentation policy; PHIPA; IPC guidance (January 28, 2026); Auditor General special report (May 12, 2026). |
| §3 Conceptual foundation. Verification defined; terminology table; four question sets on claim, system, verifier and prover (§3.2.2). | 3 Analytical framework and method. The four question sets applied to documents rather than to papers. | Single-case document analysis; each patient-facing commitment coded evidenced, asserted or absent (S6); the dimensions table in S5. |
| §4 Four fallacies, drawn from prior technical work. | 4 Findings. Four failure modes observed in this deployment. | F1–F4 below, each with its evidence status. |
| §5 Discussion. Report abstractions; do not let verification replace access; weigh privacy; institutional alternatives. | 5 Discussion. Governance and sociotechnical implications. | Physician attestation (Auditor General, Recommendation 6); patient-facing evidence; refusal without loss of care; a working correction channel. |
| §6 Conclusion. | 6 Conclusion and limits. | The bounded finding in S6: no claim of patient harm, only a claim about what the programme chose to measure. |
Section 4 turns each fallacy into a failure mode the record can show. The tag under each card gives its evidence status from the S6 coding. F1 and F3 carry the argument; F2 and F4 support it.
F1. Definitional capture (§4.1)
The procurer alone decided what “qualified” measures. Accuracy carried 4% of the score, bias controls 2% and domestic presence 30%, with no minimum passing score on any of them (Auditor General, Fig. 6, p. 21). None of the patient expectations listed in S5 has a row in that scoring.
evidenced carries the argument
F2. Component evidence, workflow claim (§4.2)
The time result comes from nine simulated encounters in a lab and from self-report in the field (Table 1; Figure 2), and accuracy was tested system by system during procurement (Auditor General, Fig. 7, p. 23). The claim that travels is about the whole documentation workflow, whose joining step is the physician's review, and no approved system required an attestation step (p. 24).
asserted supports the argument
F3. Delegated trust without a record (§4.3)
Trust did not disappear; it moved onto “qualified on the vendor list” and “the doctor will review” (p. 5359). The accuracy testing ran inside procurement in 2024 and reached the public through the Auditor General in May 2026, and the patient side of the evaluation rests on six interviews. At the time of deployment, a patient had nothing to inspect.
absent carries the argument
F4. Power outside the tool (§4.4)
The body answerable for the vendor list also wrote its weights, and the time a physician has for review is set by clinic workload that the testing never examined. Neither sits inside the systems that were tested. This reading is interpretive and needs the most careful wording in the paper.
interpretive supports the argument
Checks before drafting
- Bovens enters Paris et al. in §2.1 (p. 5350) as background on public accountability. Cite it there, not as part of the §3 lens.
- Keep the introduction's conflict at the level of evidence. The record shows an asymmetry in what was measured. It does not show a breakdown of patient trust or a shift of blame, so the paper should claim neither without new data.
- “Empirical” here means documentary evidence. Section 3 should state the corpus, the collection date and the coding rule, and every finding should read as a finding about the public record.
- Keep attribution in three layers: misabstraction is de Troya et al.'s framework, the four fallacies are Paris et al.'s, and F1–F4 are this project's.
- A methods source for single-case document analysis (for example Yin on case study research, or Bowen on document analysis) still has to be chosen and checked before it is cited.
Working documents
- Case study writing architecture (Google Doc): section-by-section outline, argument moves, evidence, counter-readings and open items for the pitch.
- REB prep: Reddit content analysis limitations (Google Doc): access, privacy, sampling and coding limits (L1–L7), each with its source status and the questions to bring to the REB.
Both documents are working drafts, dated 16 September 2026.
Open research questions arising from the visibility gap between provider efficiency and patient safeguards:
01 Is physician review a substantive safeguard, or does clinic time pressure turn it into a formal signature?
02 What can a patient actually find out before, during, and after an AI-assisted visit?
03 What specific technical and clinical criteria does provincial vendor qualification verify?
04 If verification only checks what we plan to look for, what unknown risks remain invisible?
Community literacy work is where you see that everyday people face complex forms, consent dialogues, and institutional notices — not abstract AI models. Teaching feeds this inquiry into how public evidence is communicated.
Working with learners
Literacy Unlimited (Pointe-Claire, QC) — adult English literacy. Volunteering beginning in 2026. Beginning 2026 · arrangement in progress
Their definition of literacy includes understanding health documentation and navigating official forms. In PIAAC 2022, 22% of Quebec adults aged 16–65 read at Level 1 or below (Fondation pour l'alphabétisation), against 19% for Canada as a whole (Statistics Canada) — the readers who meet multi-step instructions and official paperwork with the least margin.
Community engagement context; no research participants are recruited from this organisation.
Teaching practitioners
Long-term exploration of AI literacy, privacy experience, and emerging research gaps in intelligent systems:
UX For AI Reading Group — a 6-week cohort with 228 cumulative participants (bilingual EN/ZH). AI/UX book co-creation: AI Literacy Formation Journey (22 chapters, bilingual EN/ZH).
Columns: AI Literacy and Privacy Experience · Agentic UX: Research Gap Finder. Reference: Co-Creation on the PrivacyUX column.
All figures on this page were re-checked against the primary sources on 3 September 2026. Where a number is widely quoted but does not appear in the document it is attributed to, the discrepancy is stated in the text rather than smoothed over. Page-level citations to Paris, Moon and Guo are given as recorded by the author and should be read against the article's own pagination, pp. 5348–5370.
- Information retrieval & literature search: Perplexity was used for initial news discovery. Claude was employed to cross-check factual consistency and retrieve related academic citations.
- Web development & structuring: Cursor was utilized to adapt the HTML/CSS template structure and maintain cross-page linking.
- Editing & polishing: Grammarly was used for sentence fluency and grammatical correction.
- Intellectual content: All research questions, analytical matrices, theoretical framing, and core critical arguments were conceived, structured, and authored entirely by the human researcher.