CareCall — Can AI Reconstruct an API Contract From Legacy Code?
AI Experiment
Three AI coding assistants -- Claude Code, OpenAI, and Grok -- independently analyzed the same legacy healthcare-scheduling application and tried to reconstruct its API contract from scratch, without being told what it does and without seeing each other's work.
Question
Can an AI coding assistant accurately reconstruct an existing application's API surface, authentication model, and security posture purely from source-code evidence -- and where does it succeed, guess wrong, or honestly admit it doesn't know?
Approach
Each participant received the same prompt: analyze CareCall (a Supabase Edge Functions + Telnyx voice app) read-only, in a separate external session, and produce five evidence-backed deliverables -- a capability inventory, an endpoint inventory, an OpenAPI 3.1 spec, a findings report, and an evidence map linking every claim back to a file:line citation. Confidence had to be labeled CONFIRMED, INFERRED, or UNKNOWN; inventing behavior "because it would normally be expected" was explicitly disallowed. Separately, each participant also answered the same 10-question System Understanding Assessment -- covering tenant isolation, PHI handling, RBAC, webhook security, and more -- to test genuine comprehension rather than plausible-looking output.
Key Finding
All three independently confirmed CareCall's real architecture (5 Supabase Edge Functions, no traditional REST router, Telnyx-driven voice AI, clinic-scoped multi-tenancy via Postgres RLS) and converged on the same two critical security gaps: the Telnyx call-events webhook has zero signature verification, and the campaign dialer accepts any Bearer-shaped token with no role check behind it. But they diverged in revealing ways. Claude's five-artifact evidence set didn't cover the portal UI or read the AI assistant's instructions file directly, so it answered UNKNOWN on Campaign Reporting and on where the AI's personality is configured -- OpenAI and Grok both traced into the actual page components and config file and answered confidently. OpenAI alone surfaced a real, separate gap the other two missed entirely: a later migration that lets any authenticated user read every clinic's configuration row across tenants. Capability counts differed by scope choice, not accuracy -- Claude counted 30 granular capabilities across 22 operations; Grok grouped the same functionality into 6 capabilities and 12 operations.
What We Learned
All three tools produced a genuinely useful, evidence-grounded hypothesis of the API contract, not a hallucinated one. But evidence-grounded isn't the same as complete: each tool's blind spots were different, and no single run caught everything the other two did. Comparing independent runs side by side -- rather than trusting any one of them -- surfaced real findings that would otherwise have been missed, and that comparison is the actual point of this exercise.
Workstreams
The real workstreams, artifacts, and assessment responses behind this project — not just the summary above.