Knowledge Products
If you want AI to do knowledge work, then you need to build knowledge products.
For AI to do knowledge work that creates tangible value, it needs something better than a search box pointed at a heap of files.
Over the past decade most large organizations have learned how to make structured data usable. Warehouses, then lakes, then semantic layers and most recently the shift to treating each dataset as a product with an owner and a contract. A well-run business can now hand an analyst or an agent a clean and governed path to its numbers.
The unstructured half of the estate has had no such journey. Policies, contracts, case notes, incident reports, research, the recorded reasoning behind past decisions: this is where most of what an organization actually knows is held, and it mostly sits as a heap of files behind a search box nobody trusts. That was tolerable while the only consumers were people who knew where to look. It stops being tolerable the moment you want an agent to do knowledge work. It is even worse if you consider the volume of dormant intellectual property that could be utilized for the benefit of the business.
Agents raise the stakes in both directions. For the first time something can read across the whole heap and reason over it. But an agent grounding an answer in a stale or mis-scoped document does so with the same fluent confidence it brings to a correct one. The heap used to be a productivity problem. Connected to an agent it becomes a correctness and trust problem.
We know how to make a dataset consumable, so the question is: what is the equivalent move for a body of unstructured knowledge? The answer is the knowledge product.
A knowledge product is an owned, governed and versioned body of knowledge that people, systems and agents can safely ground decisions in.
Three products, one hotel
Picture the systems for running a hotel.
The property management system is an operational product (read “normal” product). It is where a booking is taken, a room assigned, a payment settled, a guest checked in. The system of record the hotel runs on. It transacts.
The data product is the occupancy model. A revenue manager asks how full last Friday was and it answers 94%, nine points up on the same week last year. It measures.
The knowledge product tells that same revenue manager why. Taylor Swift is playing the stadium across the road, the announcement moved every hotel in the city the day it landed, and the three comparable event weekends last year all ran around twenty points above baseline. It explains.
The concert could be modelled as data eventually, but in most organisations it first appears as unstructured context: a press release, an events calendar, a news feed, a note from sales. None of that belongs in the product raw. It enters when the owner has verified it, or when enough independent authoritative sources agree that confidence is earned rather than assumed. That act of curation is the work and added value of the knowledge product, and it is what lets the revenue manager, or an agent acting for them, trust the answer rather than hope about it.
The missing twin of the data product
Start from the structured side, because it is settled. A data product is what you get when a team treats a dataset the way a product team treats a product. Dehghani made this one of the founding principles of data mesh and called the data product the unit the whole architecture is built from. It has an owner, a published interface, documented quality, and consumers it is accountable to. DJ Patil framed the idea earlier still. The point has been understood for years. A dataset becomes useful at scale when somebody owns it and promises something about it.
A knowledge product is the same move applied to the messy half of the estate. Take a body of unstructured knowledge, give it an owner, wrap it in a contract, set a service level objective, and guarantee where every answer came from and how much we can trust that source. What was a folder of documents becomes something an agent can ground in and a person can trust.
That requires a mindset shift. Corporate knowledge is often managed as something to hoard, archive or search. A knowledge product treats it as something to curate for consumers: people, systems and agents with a job to do. The question is no longer “where should we store this?” but “who needs to rely on this, for what decision, and under what promise?”
Structured data is consumed by querying. You pose a precise question and get an exact answer, and when it is wrong it is wrong in a way you can point at: a figure that will not reconcile. Unstructured knowledge is consumed by grounding. You retrieve the passages that bear on a question and reason from them, and when it fails it fails softly, with an answer that reads well and rests on a paragraph superseded eighteen months ago. Each shape needs a product. They are not the same product.
The structured side took roughly a decade to make the leap, from datasets thrown over the wall to products with owners and contracts. The unstructured side is at the start of the same journey. We have been standing up vector stores for a while now, but almost always single-use ones: one corpus, one index, one application, built by whoever needed it and quietly abandoned when they moved on. That is the unstructured equivalent of the one-off dataset. It works for the job it was built for and compounds into nothing. The knowledge product is the same leap, made for knowledge and its potential future use cases.
We can apply the same proven design principles from the data product world to knowledge products, specifically following the DATSIS principles. A knowledge product is only as reliable as the discipline behind it. It must be:
- Discoverable, so consumers can find it in a central catalog.
- Addressable, reachable via unique identifiers and standardized interfaces.
- Trustworthy, with transparent quality checks, lineage and SLOs.
- Self-describing, carrying the metadata necessary for understanding.
- Interoperable, adhering to shared protocols for easy integration.
- Secure, with strictly enforced access controls.
By baking these principles into the product specification, we move beyond simple file storage to a governed asset that agents and systems can reliably depend on.
What a knowledge product is not
A knowledge product is not merely a knowledge base, although this is where the journey often starts. A knowledge base is storage. It will hold a decade of decayed wiki pages and tell you nothing about which are still true, which are authoritative, or who is accountable for their quality. Making that knowledge base useful and reusable is what sets it on the path to becoming a knowledge product. It needs an owner, a contract, metadata, freshness expectations and a way to act when the agent gets the wrong context at the wrong moment. Storage rots in silence. A product has someone whose job is to notice.
A knowledge product is not a single-use vector store. Standing one up does its one job, and the failure is not in building it but in mistaking it for a product. The index has no notion of which document is authoritative, which is retired, or who is allowed to see what. It answers every question with equal confidence and no accountability. The work of becoming a product is the curation, the metadata, and the contract that the one-off skips. It is the measuring and testing of value with potential consumers and customers-to-be.
Nor is a knowledge product sealed off from the structured side. The relationship runs both ways. Classify a stream of customer comments into positive and negative and what was a knowledge product is now feeding a data product. The two product types are twins in discipline, not strangers in practice. They feed each other constantly, and a mature estate is the one that has made that exchange deliberate instead of accidental.
What a knowledge product can be
Knowledge products are not a single thing. They range from document collections to formal models of meaning, and a mature estate will usually run several at once.
- A retrieval index over a curated document collection. The common case: keyword search, vector search and metadata filters over a deliberately chosen set of documents. Use all three rather than vectors alone, because exact tokens such as policy numbers, part codes and customer IDs are often what embeddings blur.
- Enterprise search. The same idea widened across many document collections behind one entry point. The hard parts are federation, ranking and access control: reconciling results and permissions from systems that were never designed to agree.
- A glossary or data dictionary. A product whose job is to say what a term means and which definition is canonical, so “monthly active user” has one meaning and not three. Cheap to build, high leverage.
- A taxonomy or ontology. A taxonomy classifies. An ontology adds rules about what an entity is and how it may relate to others. Before commissioning a new one, look for the ones the organisation has already built and shelved.
- A knowledge graph. Entities, typed relationships and a way to ask about both. This is where structured and unstructured knowledge join. Entity resolution alone often earns its keep before any full ontology does.
- A domain AI model. A fine-tuned model, classifier, reranker, extraction model or specialist agent can embody knowledge about how to classify cases, recognise risk or apply policy. If other systems depend on that behaviour, it needs ownership, versioning, evaluation, scope and a retirement path.
The retrieval contract
For structured data the contract is the semantic layer, the thing that fixes what revenue means so an agent chooses a definition rather than guessing one from column names. For unstructured knowledge the equivalent is the retrieval contract. Retrieval is approximate by its nature. The contract is what makes it precise about which source, which version, and as of when.
The contract is also a product interface. It has to work for both machine and human consumption: how systems query it, how people inspect it, how citations are followed, how uncertainty is shown, and what job the consumer is trying to get done.
A retrieval contract sets out:
- How documents are split into retrievable units. A contract and a chat transcript do not chunk the same way, and the choice trades precision against context.
- Which embedding model built the index, and the discipline for changing it. A new model is a migration. Stand up a parallel index, evaluate it, then cut over.
- What metadata travels with each unit. Authority, effective dates, what supersedes what, when it was last reviewed.
- What a citation points at. A document, a section, a passage.
- What the product promises. Retrieval precision, citation coverage, a bound on how stale an answer may be. This is impossible to define generically but very important to do specifically for each product.
Encode it and enforce it and the heap becomes a governed product. Leave it implicit and you have storage with a vector index bolted on.
Ownership and governance
Most governance for a knowledge product belongs with the product. The platform can provide tooling centrally, but the decisions have to be made locally: which sources are authoritative, who may use them, how freshness is defined, and what happens when the product fails.
The first control is access. Permissions on the source must follow the content into the product and be enforced wherever that knowledge is consumed: search, reporting, automation, human workflows or agentic systems. Access should also be explicit for non-human consumers. A service, workflow or agent using the product needs scoped, auditable authority of its own, not accidental reach inherited from whatever integration happened to be convenient. Skip this and the knowledge product becomes a breach waiting to happen.
The second is authority. The product must know which version of a thing is current, which is retired, and who has the right to decide. Knowledge has a lifetime: policies expire, guidance is superseded, product claims change, and local exceptions stop being valid. Without explicit review and QA, a knowledge product slowly becomes a better-organised archive.
The third is lineage. Every answer should trace to the source, version, timestamp and retrieval that produced it, so the product can tell whether the agent is grounded in something current or something overdue for review. A knowledge product is not fixed once it is built. The policy behind a refund changes, the chemical lookup gets a new entry, the taxonomy picks up a node it did not have last quarter. Each change has a downstream impact on whatever was grounded in the old version, just as a schema change in a data product can break whatever queried it. A mature knowledge product tracks those dependencies and knows when a change should trigger review elsewhere, not just how to serve the latest version. This is the line between a system you can defend to a regulator and one you can only demo.
Scoping an agent’s dependencies
The range of knowledge products is most useful when it is applied backwards from a use case. Take what an agent is meant to do and ask what it depends on. An anti-money-laundering check on a split payment might need an operational product to confirm the transaction happened, a data product to confirm the threshold, and a knowledge product to confirm the current compliance policy and why it applies. Once a use case is broken down this way, you can audit what already exists against what is needed. Some dependencies will already be products. Some will be a folder no one has touched in two years. The gap between the two is the actual scope of the agent build, and it is a smaller and more honest scope than “connect an agent to everything we have.”
This is also how the estate becomes tractable. Instead of connecting an agent to “all company knowledge”, you slice the problem into domain-owned knowledge products with clear consumers and explicit links to neighbouring products. Done well, that reduces cognitive load, accelerates reuse and starts to form something like a digital twin of the organisation’s working knowledge.
Why build them
The argument for building knowledge products, rather than creating a generic assistant pointed at your files, is trust. This is an engineered property of the knowledge product. If the sources behind an answer are owned, curated, versioned and fresh, the user does not have to guess whether the agent found the right page. The product carries that burden. It tells the agent what is authoritative, what is retired, what is overdue for review, and where each answer came from. A knowledge product does not make an agent correct by itself. It makes the evidence current, inspectable and governed; the agent (or human) still has to reason correctly from it.
The second argument is that you keep the feedback loop. An agent working against your knowledge products is the best instrumentation you have ever had on their quality. The questions it fails to answer become a ranked backlog of what is missing, mis-tagged or out of date. The sources nobody ever retrieves are the ones to retire. A drop in retrieval precision after a change tells you the change broke something before a user does.
Like any product, a knowledge product has a lifecycle. It is built, enriched, tested, contextualised, released, monitored and retired. That lifecycle is increasingly AI-supported and AI-supportive: AI can help classify, link, summarise and test knowledge, while the resulting product gives AI something governed to depend on. This is where a folder becomes a domain-aligned asset, closer to “our proprietary pricing engine” than “some documents in SharePoint.” As a side effect, the actual current benefit for an organisation can also be framed as follows: a knowledge product for an agent is a thin slice to evolve an AI-enabled product development lifecycle.
Capture that signal and the estate improves as a by-product of being used. Each interaction leaves the next one a little better off. Organizations that wire up this loop pull steadily away from the ones that treat their knowledge as a static pile, and the gap compounds.
In short
For AI to do knowledge work, it needs knowledge products: document collections, glossaries, graphs, embedding models and specialist AI models with owners, contracts, permissions, authority, freshness, lineage and evidence of quality. A data product stops an agent inventing the numbers, a knowledge product stops it silently trusting the wrong source. Product thinking is what turns either one from a pile of data into something people and systems can depend on. Customer centricity is what helps you find out what part of the organisation’s IP might become a knowledge product.
Acknowledgements
Big thanks to many of my Thoughtworks colleagues for their additions, questions and downright disagreement! Those who have made additions have been added as co-authors, but I would like to mention the following (in no particular order) for their discussions, comments and suggestions: Fabian Nonnenmacher, Moritz Wilke, Danilo Sato, Paola Attadio and Tim Harrison.