This week: Mass General and UCLA publish the safety architecture behind an opioid-recovery chatbot, a Madrid surgeon simulates an appendectomy with GPT-6 Astra, and the directory has its biggest week yet, 15 new clinician-builders.
Top signal
What a chatbot should do when the patient on the other end is in crisis. Researchers at Massachusetts General Hospital and UCLA, publishing in npj Digital Medicine on 22 September, describe the safety architecture behind Suzy, a chatbot already piloted with patients receiving medication treatment for opioid use disorder at a Boston addiction clinic. Every message first hits a "Safety Router": a classifier trained on 200 simulated messages that four clinical experts ranked into three tiers, from clear risk (their example: "I am thinking about ending my life") through ambiguous distress to no risk at all. A Tier 1 message triggers an immediate safety disclaimer, and the authors are blunt about what should happen next: "a notification to a clinician, care team member, or centralized monitoring service is essential", because AI monitoring with no human in the loop has already been shown not to be enough. The system runs on a Business Associate Agreement and a zero-retention deal with OpenAI. It is one pilot at one clinic, not a validated standard, and the paper says so; still, it is the most concrete published answer yet to a question every clinician shipping a chatbot eventually has to face. No regulator published anything comparable this week. npj Digital Medicine
Shipped this week
- Wakhile Keiren Mavimbela: the nursing student at the University of Eswatini built PhiliGo, an offline-first web app with four treatment-adherence games and a caregiver dashboard for children on tuberculosis treatment, in vanilla HTML, CSS and JavaScript, pushed this week. No stars yet and no clinical evaluation, but a TB adherence tool built for children, by a nurse, in a country carrying a heavy TB burden, is exactly the kind of ship this list exists to catch early. GitHub
- Julio Mayol: the Madrid professor of surgery posted his own working laparoscopic appendectomy simulator, built with OpenAI's GPT-6 Astra: "just built a digital tool to simulate a lap appy", he wrote, framing it as bridging the gap between AI and surgical training. No link to the tool itself and no clinical validation yet, so read it as a demonstration of what a senior surgeon can now prototype solo, not a ready teaching resource. X
- Sohrab Arora, MD: the Fort Worth urologic oncologist shipped Urology Oral Boards Coach, a website and iOS app built on physician-authored case frameworks rather than open-ended chat: "this app is not just a wrapper around a chatbot", he wrote, pointing to more than 300 authored oral-board cases run on a deterministic rubric, over 100 OSCE scenarios and a voice mode for practising on the commute. No independent testing beyond his own use yet. X
- Max Brzezicki: the Oxford neurosurgery clinical fellow (MB ChB, DPhil) published the full analysis package behind his admission-glucose and thrombectomy-outcome study: a reproducible Jupyter pipeline bundled with a provenance manifest and a script that checks every number in the manuscript against the analysis output, so the paper and the code cannot silently drift apart. Zero stars so far; the point is not popularity but that anyone can re-run his numbers. GitHub
Build safely
- The number that actually measures AI safety in radiology: a team at Bern University Hospital's neuroradiology department measured what matters more than adoption: how often radiologists actually reviewed and explicitly accepted or rejected the AI's output before acting on it, rather than clicking through. Before a vendor-neutral orchestration platform and a structured change programme (training, workflow standardisation, feedback, governance), that verification rate was 13.9 per cent (318 of 2,287 outputs); afterwards, 61.1 per cent (1,778 of 2,910). Published 22 September in npj Digital Medicine. Vendor content: the orchestration platform is Bayer's Calantic, an early-adopter partner that supported the implementation, though the authors state it had no role in study design or analysis. npj Digital Medicine
- Who is accountable when the commercial AI gets it wrong: University of Southampton researchers argue the accountability gap in digital health is not mainly technical: AI tools trained on data that under-represents ethnic minorities, rural patients, people with disabilities and people experiencing homelessness reproduce the access gaps already built into who gets recorded in the first place, and NHS procurement has no real mechanism to check for it. Their ask: transparent subgroup-performance reporting, representative datasets, and monitoring that continues after deployment, not only before. Published 17 September in the Journal of Medical Internet Research. JMIR
Tools & guides
- A one-click PHI scrubber before you paste into AI: Arman, an infectious diseases physician, built ScrubBoard on a whim: an app that strips protected health information from text before you paste a case into a general-purpose AI assistant. "Still needs testing though! Very early", he wrote when he shared it. No clinical validation and no claim it catches everything, so treat it as a first pass, not a compliance guarantee, before you paste real patient text anywhere. GitHub
Events
- 3 Oct (Antwerp): Care & Code Clinical Build Day: doctors, nurses, pharmacists and physiotherapists vibe-code a working AI care tool in one day, tickets on sale now. Disclosure: this is our own event. careandcode.be
- 26 Sep (Vienna, next application batch): AIM Austria's healthcare hackathon (29-30 October, IBM's Vienna headquarters) closes its next application batch tomorrow, with two batches still to follow: three tracks on real partner use cases, over EUR 15,000 in cash and credits, and a follow-on venture track with investor introductions for standout teams. healthcare-hackathon.eu
New in the directory
- Chisom Rutherford, MBBS (Nigeria): built housejob.ng, a clinical decision-support platform she says is in use by more than 1,300 doctors, alongside a personal machine-learning notebook on asthma classification on her own GitHub. housejob.ng
- The directory now counts 70 clinician-builders with verified, evidence-graded ships, its biggest week yet: 15 joined, no removals, so the net is 15 more than last week. Several were already familiar from earlier issues (Dan Heslinga, Youssef Ahmad and Shardul Dhande among them) formally joining now that their entries are verified. Browse it, forward this to a colleague who ships, and reply if you know someone who belongs in it. cliniciansthatcode.com