<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>Jeonghwan Kim — Writing</title>
    <link>https://jeonghwankim.dev/writing</link>
    <atom:link href="https://jeonghwankim.dev/feed.xml" rel="self" type="application/rss+xml"/>
    <description>I build and ship software end-to-end — from system design to deployment. Full-stack web, native iOS/watchOS, and real-time systems, across AI, audio, mental health, legal tech, and neuroscience research.</description>
    <language>en</language>
    <item>
      <title>The Model Wasn&apos;t the Problem. The Scoreboard Was.</title>
      <link>https://jeonghwankim.dev/writing/the-model-wasnt-the-problem</link>
      <guid isPermaLink="true">https://jeonghwankim.dev/writing/the-model-wasnt-the-problem</guid>
      <pubDate>Thu, 06 Aug 2026 00:00:00 GMT</pubDate>
      <description>My cross-validation reported 0.61. The pooled out-of-fold number was 0.52. Both were correct, and the gap had nothing to do with the model — it was five folds scoring on five incompatible scales. Part 5 of The Differential.</description>
    </item>
    <item>
      <title>From Case Reports to CT Scans</title>
      <link>https://jeonghwankim.dev/writing/from-case-reports-to-ct-scans</link>
      <guid isPermaLink="true">https://jeonghwankim.dev/writing/from-case-reports-to-ct-scans</guid>
      <pubDate>Tue, 04 Aug 2026 00:00:00 GMT</pubDate>
      <description>The text benchmark left me arguing about what counts as a finding. So I moved to pancreatic CT, where a wrong answer is just wrong. The first run scored 0.52 AUROC — and the way it failed was worth more than the score. Part 4 of The Differential.</description>
    </item>
    <item>
      <title>The Model Didn&apos;t Hallucinate. It Still Failed.</title>
      <link>https://jeonghwankim.dev/writing/the-model-didnt-hallucinate</link>
      <guid isPermaLink="true">https://jeonghwankim.dev/writing/the-model-didnt-hallucinate</guid>
      <pubDate>Sat, 01 Aug 2026 00:00:00 GMT</pubDate>
      <description>50 manual extraction runs across two clinical NLP benchmarks. Source grounding was 100% and the unsupported-claim rate was zero — but the evidence units, entity boundaries, and temporal relations fell apart. Part 3 of The Differential.</description>
    </item>
    <item>
      <title>When the Evidence Changes, Does the Model Really Update?</title>
      <link>https://jeonghwankim.dev/writing/when-evidence-changes</link>
      <guid isPermaLink="true">https://jeonghwankim.dev/writing/when-evidence-changes</guid>
      <pubDate>Wed, 29 Jul 2026 00:00:00 GMT</pubDate>
      <description>I repeated every perturbation condition three times. The leading diagnosis was highly reproducible; the ordering and scoring below it was not — which killed my original product hypothesis. Part 2 of The Differential.</description>
    </item>
    <item>
      <title>Getting the Diagnosis Right Is Not the Same as Reasoning Reliably</title>
      <link>https://jeonghwankim.dev/writing/getting-the-diagnosis-right</link>
      <guid isPermaLink="true">https://jeonghwankim.dev/writing/getting-the-diagnosis-right</guid>
      <pubDate>Wed, 29 Jul 2026 00:00:00 GMT</pubDate>
      <description>A frontier model ranked the correct diagnosis first on a masked neuromuscular case. Then I started removing evidence — and its confidence didn&apos;t move the way it should. Part 1 of The Differential.</description>
    </item>
    <item>
      <title>A self-audit of a production SaaS</title>
      <link>https://jeonghwankim.dev/writing/a-self-audit-of-a-production-saas</link>
      <guid isPermaLink="true">https://jeonghwankim.dev/writing/a-self-audit-of-a-production-saas</guid>
      <pubDate>Sun, 19 Apr 2026 00:00:00 GMT</pubDate>
      <description>I sat down for a day and attacked my own app. Six categories of finding — and what I&apos;d build into the next project from day one.</description>
    </item>
  </channel>
</rss>
