Why AI Responses Need Structure, Not Just Markdown

8/26/2026

Conversational text versus structured response data

We recently fixed something almost embarrassingly small: an AI assistant's replies were showing up full of stray asterisks and pound signs instead of real formatting. Headings looked like raw Markdown syntax. Bullet points looked like dashes and dots instead of a list. The model was writing perfectly good Markdown; the interface just wasn't rendering it. It was a quick fix, and also a useful reminder of something bigger: text formatting is the easy part of the problem. The harder part is that a conversational reply and the useful thing the user actually wanted out of it are not the same shape of data, and treating them as one flat string of prose caps how useful an assistant interface can be.

A good response to "which approach should I use here?" isn't just an essay. It's an answer, plus maybe a one-line summary, a couple of sources, a short list of what's still needed, and a next step or two the user can act on. Cramming all of that into paragraph form and hoping the user reads carefully is a UI failure dressed up as writing style.

Quick summaries and next actions

Long, well-reasoned answers are often exactly what's needed -- and also exactly the kind of thing a busy person skims past looking for the actual outcome. A short, separately rendered summary alongside the full explanation solves both problems at once: the reasoning stays available for whoever wants it, and the headline doesn't get lost inside it.

The same goes for next steps. "Here's what I'd do" is more useful as a small, scannable list than as the last sentence of paragraph four. None of this replaces the conversational reply -- it sits alongside it, pulling out the parts that deserve to be seen at a glance rather than read for.

Citations and generated documents

Once an assistant is doing real research -- searching internal knowledge, checking prior work, pulling in project context -- its claims should come with receipts. Not links invented after the fact to look credible, but citations traceable to what the assistant actually retrieved and used while forming that specific answer. If a source wasn't part of the evidence behind a claim, it shouldn't be presented as if it were.

The same discipline applies to documents. An assistant that says "I've created a plan for you" needs that plan to actually exist as a real, openable thing -- not a description of a document that never got written. A proposed document and a produced one are different states, and the interface should never blur them into looking the same.

Trusted application navigation

An assistant that understands an application well enough to talk about it should also be able to point at it -- "here's the relevant page," "here's where that lives." But a language model proposing a URL is not the same as a validated, safe link. The model can identify intent -- which project, which article, which setting -- while the application itself resolves that intent into an actual, permission-checked destination. A link a user can click should never be something the model typed out freely; it should be something the surrounding system built and vouched for.

Artifact collections

Over a longer conversation, useful things accumulate: documents produced, sources cited, decisions made, loose ends still open. Buried across dozens of messages, that accumulation is only as accessible as someone's patience for scrolling back. Collected into one place, it becomes something closer to a working file for the conversation -- the durable outputs, distinct from the back-and-forth that produced them.

That distinction matters more than it sounds: routine "here's a link to that page" navigation isn't an artifact, and treating everything the assistant ever points at as equally durable output just recreates the clutter problem one level up. The collection is only useful if it stays a curated view of what actually matters.

Graceful fallback when structured generation fails

Structure is a layer added on top of a working conversational reply, not a replacement for one -- and it has to be allowed to fail without taking the reply down with it. Getting a model to reliably return well-formed structured output alongside natural conversation is a genuinely hard problem, and any implementation of it will occasionally produce something malformed, incomplete, or from a provider that doesn't support it well at all.

The right response to that failure is quiet degradation: show the plain conversational answer, drop the structured extras for that one response, and move on. A user should never see raw formatting artifacts or lose the answer entirely because the nicer version of it didn't come together. Structure should make good answers better. It should never be a single point of failure standing between a user and a working reply.