Birth of AI models

How an AI Model Is Born — and Why Where It Lives Matters

9/9/2026

In the last piece, I made the case that knowledge is the resource every consequential era has run on, and that the AI era's version of that resource — structured, retrievable knowledge — should behave more like a commodity than a luxury good. This piece is the "how" behind that. If knowledge is the grain, this is the mill: how raw material becomes a working model, and where that model actually lives once it's built.

Birth: from raw text to a working model

A model doesn't start as a system that "knows" things — it starts as a very large pile of undifferentiated text, code, and data, fed through a training process that teaches it statistical patterns of language and reasoning. That base model is a generalist: broad, capable, but not grounded in any one organization's facts, and not up to date past whenever its training data was collected.

That gap — general capability without current, specific, verifiable knowledge — is exactly where the rest of this stack comes in.

Fine-tuning: teaching the generalist a specialty

Before a model ever gets to retrieval or tools, there's often a shaping step: fine-tuning. A base model knows language broadly; fine-tuning nudges it toward a narrower skill or domain by training it further on a smaller, more targeted dataset — a customer support tone, a legal domain's vocabulary, a company's own way of answering.

Full fine-tuning — updating all of a model's parameters — is expensive and heavy, which is part of why lighter approaches have taken over for most real-world use. QLoRA (Quantized Low-Rank Adaptation) is one of the more important of these: it freezes the bulk of the original model, compresses it to save memory, and trains only a small set of added parameters that capture the new behavior. The result is a specialization step that's achievable on a fraction of the hardware full fine-tuning would need.

Incremental fine-tuning takes this further — rather than retraining from scratch every time new correction data shows up, a model is updated in small, repeated passes as new examples accumulate. This is closer to how the KB Sandbox framework I've written about elsewhere treats it: human corrections at the point of use become new training data over time, and fine-tuning is a periodic, deliberate step layered on top of a system that's already grounded through RAG — not a replacement for grounding, but a way to bake in behavior that retrieval alone can't fix.

RAG: giving the model something to stand on

Retrieval-Augmented Generation (RAG) is the bridge between a general-purpose model and a body of specific, current knowledge. Instead of relying purely on what the model memorized during training, RAG lets it look something up before it answers: a query goes out to a knowledge store, relevant passages come back, and the model reasons over those passages instead of guessing from memory.

This matters for two reasons. First, it's how you keep a model current without retraining it — update the knowledge store, not the model. Second, it's how you make a model's answers checkable: if the answer is grounded in a retrieved passage, you can trace it back to a source instead of trusting a black box.

Wikis and LLM knowledge bases: the raw material, organized

RAG is only as good as what it retrieves from. This is where wikis and structured knowledge bases do the quiet, unglamorous work that makes everything else possible — curated, cross-referenced, continuously corrected bodies of knowledge that give a retrieval system something worth retrieving.

This is the direct continuation of the historical pattern from the last post: noise becomes organized knowledge, and organized knowledge becomes leverage. A wiki that's actually maintained — scrutinized, updated, disputes resolved — is the modern equivalent of a well-kept manual. A wiki that's stale or unmoderated just adds noise back into a system that was supposed to remove it.

Agent tools: knowledge that acts

The next layer up is agency — giving a model tools, not just knowledge. An agent isn't just retrieving and answering; it's deciding when to search the internet, when to query a database, when to call an API, and when to hand a decision back to a human. This is what turns a knowledge system from something you consult into something that does work on your behalf.

The important design discipline here is constraint. An agent with unrestricted tool access is a liability, not a feature. The systems worth trusting define explicitly what an agent can search, what it can act on, and where a human has to sign off before anything irreversible happens.

Distribution: where the intelligence actually lives

The last piece of the puzzle is physical, not conceptual: where does all of this run?

The default answer for the last decade has been central cloud — large, centralized data centers where the heaviest models live, reachable over the internet. This is where you get the most raw capability, but it comes with latency, connectivity dependence, and centralized cost and control.

The emerging alternative is the intelligent edge — pushing inference, retrieval, and even smaller fine-tuned models closer to where the data and the decision actually happen, rather than routing everything back to a central server. This matters most when latency is unacceptable (a factory floor, a hospital device, a vehicle), when connectivity can't be guaranteed, or when data has strong reasons not to leave its origin (privacy, sovereignty, compliance).

The realistic shape of the next few years isn't "cloud vs. edge" — it's a split: heavy reasoning and large-scale training stay centralized, while retrieval, lightweight inference, and time-sensitive decisions move closer to the edge. The knowledge layer (the wikis, the vector stores) increasingly needs to be distributable too — not just one central copy, but synchronized, current copies wherever the decision is being made.

Why this connects back to the commodity argument

Every layer in this stack — the base model, the retrieval system, the knowledge base, the agent, the compute it runs on — is a place where cost can either be concentrated (scarcity pricing) or spread out (commodity pricing). The birth of a model is expensive; that part is real and will likely stay concentrated for a while. But the knowledge layer underneath it — the wikis, the curated sources, the retrieval infrastructure — doesn't have the same structural excuse for scarcity. That's the layer where the wheat-and-rice argument applies most directly, and it's the layer this series will keep coming back to.