This week: a University of Florida team shows five AI evidence-search tools miss most of the literature on any single query, an Oxford resident ships a free ward handover app, and a Care & Code build day starts tomorrow in Antwerp.
Top signal
Ask an AI tool to find the evidence and it will, usually, miss most of it. Researchers at the University of Florida, publishing in npj Digital Medicine on 29 September, tested five tools clinicians increasingly use to search the literature (Consensus, Ai2 Paper Finder, ChatGPT, Gemini and Claude) against a non-public, hand-assembled gold-standard corpus, so the answers could not have leaked into training data. Across 15 query formulations, the median share of relevant evidence a single formulation retrieved ranged from 7.2 to 42.2 per cent depending on the platform. Pooling every formulation helped, to between 45.8 and 72.3 per cent, but even then 12.0 per cent of the evidence was never retrieved by any platform, and conference proceedings fared far worse than journal articles (38.9 per cent never found, against 4.6 per cent). For the largest evidence category, the chance that a single query returned nothing relevant at all ran from 47 to 80 per cent, and one platform showed a marked gap before 2016. The authors' conclusion is the one every clinician building a literature tool should take: evaluate retrieval in your own domain before you trust the output. It is a single study of five tools at one point in time, and the tools will have changed by the time you read it. npj Digital Medicine
Shipped this week
- Khalid Shamiyah, MD: the Oxford obstetrics and gynaecology resident built Handover, a free iPhone app for the ward, after "8pm. 30 patients. 50 jobs. One crumpled sheet of paper." Jobs are captured by ward and bed, reminders can be set, and a colleague takes over your list from a QR code or a link at the end of shift. Everything stays on the phone, with no account, no server and no cloud sync, and the App Store listing tells users to enter non-identifying shorthand only and says it does not replace verbal handover. No clinical evaluation yet, and it is iPhone only for now. X
- Shruti Malik, MBBS: the Omaha-based physician and certified pharmacy technician says she built MedTools.net, an AI tool hub for healthcare professionals, entirely herself: "Idea to launch. Me." The site lists prescription cost savings, a drug interaction checker, exam prep, skin lesion analysis and a pathology atlas. Her own post is the only source for who built it, and we have not seen any validation of the clinical tools, so treat the interaction and skin-lesion features as unproven until someone tests them. X
Build safely
- Can ChatGPT do the risk-of-bias check on a trial?: a team led from the pharmacy department of Fu Jen Catholic University Hospital in Taiwan gave ChatGPT 29 randomised trials taken from published Cochrane reviews and asked it to apply the Cochrane Risk of Bias 2 tool, twice per trial, Domain-level accuracy against the Cochrane authors came out at 73.1 per cent on the first pass and 75.9 per cent on the second; sensitivity for spotting a high risk of bias fell from 61.4 to 53.4 per cent between passes, so the model leaned towards missing problems rather than inventing them. Agreement between the two passes was 89.0 per cent overall, but a Cohen kappa of just 0.39 in the second domain. The authors' verdict is an exploratory one: useful as a first screener, never without expert oversight. Published 25 September in the Journal of Medical Internet Research. JMIR
- Fine-tune, retrieve, or both? The evidence is thin: a systematic review in the Journal of Medical Internet Research (17 September) pooled 35 studies, published between 2024 and 2026, that adapted a language model for clinical decision support: 17 used retrieval-augmented generation, 7 fine-tuning and 11 a hybrid. Its pragmatic reading is fine-tuning for narrow classification, retrieval for guideline-grounded reasoning and hybrids for complex multimodal work, with retrieval lifting guideline-adherence tasks from 71.1 to 92.1 per cent and from 78.9 to 94.7 per cent in two studies. The caution matters more for builders: 25 of the 35 studies were judged at high risk of bias, 9 unclear and only 1 low, mostly for missing external validation and calibration, and the benefits of retrieval were inconsistent on larger reasoning models. JMIR
Events
- 3 Oct (Antwerp): Care & Code Clinical Build Day is tomorrow: doctors, nurses, pharmacists and physiotherapists vibe-code a working AI care tool in one day. Disclosure: this is our own event. careandcode.be
- 3 Oct (Vienna, application batch 5 of 6): AIM Austria's healthcare hackathon (29 and 30 October, IBM's Vienna headquarters) closes its fifth application batch tomorrow; the sixth and final one closes on 10 October. Three tracks (clinical documentation, health data interoperability and the European Health Data Space) with EUR 5,000 each, and remote teams can join a separate community track, though the cash prizes are for on-site teams only. healthcare-hackathon.eu
New in the directory
- Xus Badia, MD (Spain): a Barcelona colorectal surgeon whose bio reads "Building software between surgeries", with four clinical and academic tools live on their own sites: Trialinx, a study-design and data-capture platform with a free plan, Chronosurg, a surgical registry with 12 clinical modules, LIR-AEC, a resident's logbook, and Relaylit, which emails relevant new papers. No independent evaluation of the clinical tools that we have seen. badia.me
- The directory now counts 70 clinician-builders with verified, evidence-graded ships, the same as last week: nobody joined or left. Five more people we featured earlier in Shipped sections now have verified entries: B Gurubasavaraj, Leonardo Pegollo, Julio Mayol, Max Brzezicki and Wakhile Keiren Mavimbela. Browse it, forward this to a colleague who ships, and reply if you know someone who belongs in it. cliniciansthatcode.com