Back to insights
Field Notes·9 min read

Your RAG Index Needs Schema Discipline, or It Lies

Your RAG Index Needs Schema Discipline, or It Lies

A database schema change without a migration plan gets you fired. A RAG re-embed without one gets you a demo that still works and a production system that quietly started answering worse. The knowledge base behind a RAG system is a data store with a schema, the embedding model, the chunking, the metadata contract, and it deserves the same versioning, migration, and rollback discipline you'd never skip on Postgres. Most teams treat it like a folder they occasionally re-run a script over.

By Daniel Usvyat · Principal, USQRD

Share

The Re-Embed That Changed Everything and Errored Nothing

Swapping the embedding model is the move that looks free and isn't. A newer model scores better on MTEB, someone re-embeds the corpus over a weekend, and Monday's answers are subtly different. Not broken — different. Queries that used to surface the right passage now surface a plausible neighbour. Nothing throws. No alert fires. The only signal is that answer quality drifts, and by the time support tickets reflect it you've lost the causal thread.

The reason this fails quietly: the embedding model defines the geometry of your vector space. Change the model and every distance changes. A cosine similarity of 0.82 under the old model and 0.82 under the new one are not the same relationship. Your `top_k` cutoff and your similarity threshold were tuned against the old geometry, and so were your reranker's assumptions. Re-embed without re-tuning and you've shipped a different retrieval system wearing the old one's config.

We treat the embedding model as part of the schema, not a hyperparameter you can freshen up in place. A model change is a migration. It gets a new index, a new version, and an eval gate before it goes anywhere near production traffic. The teams that skip this are the ones in our RAG audits wondering why the demo held and prod slid.

A cosine similarity of 0.82 under the old model and 0.82 under the new one are not the same relationship.

Orphaned Chunks Answer With Dead Facts

Partial reindexing is where the really nasty failures live. A doc changes, someone reindexes just that document, and the old chunks from the previous version never get deleted. Now the index holds two generations of the same content. Retrieval doesn't know one is stale, both are valid vectors with valid metadata, and the dead one wins the similarity contest often enough to matter.

We've seen a corpus where a policy was updated, the new version was indexed, and the superseded chunks stayed in place because the delete step keyed on a document ID that had changed. The agent cited the old policy, confidently, with a real source link. That's not a hallucination — the model faithfully repeated a chunk that should not have existed. This overlaps with how contradictory and outdated corpora make agents cite the wrong thing: the retrieval works perfectly, the store is corrupt.

The fix is boring and it's the same one databases learned decades ago. Reindexing is not an upsert into a mutable pile of vectors. Every chunk carries a version and a source-document fingerprint, and a reindex is transactional: the old generation is fully removed or fully shadowed before the new one serves. If you can't atomically swap a document's chunk set, you will orphan chunks eventually.

  • →Key deletes on a content fingerprint, not a mutable document ID or file path.
  • →Stamp every chunk with an index version and a source-doc hash so you can prove which generation served an answer.
  • →Never partial-reindex into a live index in place, write to a shadow, verify chunk counts, then swap.
  • →Assert on total chunk count and per-document chunk count after every reindex; a drift is a bug, not noise.

Version Indexes Like Immutable Artifacts

The pattern that makes all of this tractable: an index is an immutable, versioned artifact, and production points at one via an alias. You don't edit `prod-index`. You build `index-v47`, validate it, and repoint the alias. Rollback is repointing the alias back to `v46`. This is exactly how we think about prompts as versioned artifacts rather than editable dashboard config — the same discipline, applied to the retrieval store.

The version has to capture everything that affects retrieval behaviour: the embedding model and its exact version, the chunking strategy and its parameters, the corpus snapshot. Change any of those and you've got a new index version. When someone asks why retrieval changed on the 12th, you want to answer with a diff between two named artifacts. What you don't want is to shrug and run git blame on an ETL script.

Most vector databases support aliases or index swapping natively, pgvector via a table swap, the managed stores via alias APIs. The mechanism is cheap. The discipline of never mutating the live index is what people skip because in-place upserts feel faster. They are faster right up until the day you can't explain what your system is retrieving.

Eval the Candidate Index Before You Promote It

You do not promote an index because the re-embed job finished. You promote it because it passed a retrieval eval against the version it's replacing. And the eval that matters first is retrieval-only, did the right passage make it into the candidate set, not end-to-end answer quality. Grading answers while ignoring whether the right passage was even retrieved tells you nothing about the index change specifically.

Build a labelled set of queries with known-correct source chunks and run it against both the current and candidate index. Look at recall@k and MRR side by side. A new embedding model that improves aggregate recall by 3% but tanks a specific query class, say, anything with product codes or dates, is a regression you want to see before your users do, not after. This is where a retrieval eval use earns its keep as the actual deliverable.

One thing the eval won't catch on its own: slow drift after promotion. A candidate can pass every gate and then rot as the corpus grows and query distribution shifts, the same way agents pass evals and quietly degrade. So the eval isn't a one-time gate. It's a regression suite you re-run on a schedule and on every index build, with the golden set refreshed from real production queries before it goes stale.

You do not promote an index because the re-embed job finished. You promote it because it passed a retrieval eval.

The Operational Cost Nobody Budgets For

Here's the part that gets left out of the estimate. A full re-embed of a large corpus is a compute job with a real bill and a real duration, and you'll run it more often than you think — every embedding model upgrade, every chunking change, every time you decide the metadata schema was wrong. During that window you're maintaining two indexes and paying for both. If you want zero-downtime migrations, you need the storage and the pipeline to run the shadow build in parallel.

Then there's the incremental-versus-full decision, and it's genuinely hard. Incremental reindexing is cheap but accumulates skew — orphaned chunks, inconsistent chunking across generations, metadata that drifted as your schema evolved. Full reindexing is clean but expensive and slow. In our engagements the honest answer is you do incremental for content updates and force a full rebuild whenever the schema changes, and you accept that 'the schema' includes the embedding model. Teams that try to incrementally migrate across an embedding-model change end up with a vector space that's half one geometry and half another, which retrieves worse than either.

None of this is exotic engineering. It's the migration discipline every backend team already has, applied to a data store that happens to hold vectors. The reason it gets skipped is that a RAG index feels like a cache you can always rebuild, so people treat rebuilds as casual. They aren't casual — a rebuild is a schema migration, and an unversioned migration you can't roll back is how you end up debugging retrieval quality with no way to compare against last week.

What's Still Hard, and Where to Start

The genuinely unsolved part is cross-model eval comparability. When you change embedding models, your similarity scores aren't comparable across versions, so any threshold-based logic needs re-tuning per index — and there's no clean automated way to know your new thresholds are as well-calibrated as the old ones without a solid labelled set. If your eval data is thin, a model migration is partly a leap of faith, which is exactly why the labelled retrieval set is the first thing worth building.

Start with three things this week. Stamp every chunk with an index version and a source-doc fingerprint so you can at least answer what served a given answer. Move production behind an alias so promotion and rollback are pointer swaps, not rebuilds. And write a small retrieval-only eval, even 50 labelled queries, that you run against any candidate index before it serves traffic. That's enough to turn a silent re-embed disaster into a diff you can see before it ships.

Frequently asked questions

Do I need to re-embed my entire corpus when I change embedding models?

Yes. Embedding models define the geometry of your vector space, so mixing vectors from two models in one index retrieves worse than either alone. Treat a model change as a full rebuild into a new versioned index, not an incremental update.

How do I roll back a RAG index if a new version retrieves worse?

Keep each index as an immutable versioned artifact and point production at it through an alias. Rollback is repointing the alias to the previous version — which only works if you didn't mutate the old index in place.

What causes a RAG system to cite outdated or dead facts?

Usually orphaned chunks from partial reindexing: the new version of a document gets indexed but the old chunks are never deleted, so both generations compete in retrieval and the stale one sometimes wins. Key your delete step on a content fingerprint, not a document ID that can change.

Should I run a full or incremental reindex?

Incremental for routine content updates, full whenever the schema changes — and the schema includes the embedding model, chunking strategy, and metadata contract. Incremental migrations across a schema change accumulate skew that retrieves badly and is hard to debug.

Free resource

Take the Operational Bottleneck Audit

Our Bottleneck Audit maps where your retrieval pipeline silently degrades — from index migrations to orphaned chunks — before it costs you.

Ready to stop experimenting?

Get your RAG pipeline audited before the next re-embed

We'll find the migration, versioning, and eval gaps that make production retrieval drift silently. Senior engineers, no theatre.

Book a Discovery Call
More insights
Your RAG Index Needs Schema Discipline, or It Lies | USQRD