Back to blogData

TheDataInfrastructureYouActuallyNeedBeforeShippingAIFeatures

Julian Voss· Head of Data Platform· January 20, 2026· 8 min read

Freshness: the SLA nobody writes down

Teams frequently ship an AI feature against a data pipeline with no defined freshness guarantee, and then get surprised when the model answers using week-old inventory counts or outdated policy documents. We push every AI feature to have an explicit freshness SLA before launch, whether that is real-time streaming, hourly batch, or daily refresh, matched to how quickly the underlying source data actually changes, not to whatever the existing pipeline happens to already do.

Getting freshness wrong in either direction is expensive: too stale and the AI feature gives confidently wrong answers, too real-time when it is not needed adds infrastructure cost and complexity for no user-facing benefit. We size the SLA based on how often the answer would have actually changed if it were checked, not on what's technically achievable.

Lineage: knowing where an answer came from

When an AI feature produces a wrong or surprising answer, the first question is always where the data came from, and teams without lineage tracking spend days reconstructing that manually. We recommend tagging every record that flows into an AI pipeline with its source system, ingestion timestamp, and any transformation applied, stored alongside the data rather than in a separate document that goes stale.

This is not just a debugging convenience. In regulated industries, being able to show exactly which source record and which transformation produced a specific AI output is often a compliance requirement, and retrofitting lineage after a system is in production is dramatically more expensive than building it in from the first pipeline.

Vector stores are a database decision, not a library decision

Teams often pick a vector store based on which library had the simplest quick-start guide, then discover months later that it does not support metadata filtering, incremental updates, or the scale they actually need. Vector store choice should go through the same evaluation as any other production database: query latency under realistic load, support for filtered search alongside similarity search, backup and recovery story, and operational maturity of the team running it.

Feature stores bridge the online and offline gap

A model trained on features computed one way in a batch pipeline and served with features computed a different way in a real-time API is a classic source of silent production degradation, often called training-serving skew. A feature store that guarantees the same feature definition and computation logic in both training and serving contexts closes that gap, and it is worth the setup cost for any AI feature with more than a handful of derived features feeding it.

The minimum viable data platform

Not every AI feature needs a full modern data stack on day one. For a first launch, we typically consider the floor to be: a documented freshness SLA, basic lineage tags on ingested data, a vector store chosen with production requirements in mind rather than convenience, and a single source of truth for any feature used in both training and inference. Everything beyond that floor can usually be added incrementally as the feature proves its value.

Data PlatformVector DatabasesData EngineeringMLOps

Wanthelpshippingsomethinglikethis?