Ingestion, extraction, a resolved data model, a time series and a knowledge graph over millions of private companies. Reachable as an API and an MCP server, deployable inside your own cloud, and stopping exactly where your own judgement begins.
Your sources go in at the bottom. The ingestion layer and the time series engine turn them into one record per company, person and fund. The knowledge graph connects them. Workflows, analytics and your own agents sit on top.
The same diagram every Kruncher customer sees inside the product. The difference for a builder is that each layer is reachable on its own, so you can take the part you need and keep everything you have already built.
Each one is a hard problem on its own, and each one is where an internal build tends to stall.
A fourteen step ingestion pipeline takes documents, email, calendars, call transcripts, your CRM, your data rooms and internal registers, alongside more than twenty premium databases and public registries across several languages. Nothing is read until the engine knows which document it is looking at and which entity the content belongs to.
The structure is resolved before a single value is read, so the engine knows whether a number belongs to a vehicle, an asset, a period or a person. Every candidate value keeps the file, the page and the position it came from, and competing candidates are kept beside the preferred one with the reason it was preferred.
Roughly a thousand normalized data points per company, and each one labelled by where it came from: what the company says about itself, what the public record shows, what Kruncher estimates, and what your own team knows. Contradictions are resolved by source precedence rather than averaged, and a gap is flagged rather than filled.
Every value is dated to the period it describes rather than the day it was read, so a March deck reporting December figures produces December facts. The record is a time series, which is what makes it possible to ask what changed between any two dates and to see the trajectory rather than the last reading.
Entity resolution collapses the same company, person or fund appearing under different names across every source into one node. Around it sit typed, directional edges: investor, partner, vendor, customer, competitor, founder, executive. Millions of companies, and every edge carrying the same provenance as any other data point.
AI ready means more than a clean JSON response. It means the disambiguation, the provenance and the time dimension are already in the data, rather than something your agent has to work out.
An agent asking about a company gets one record, not twelve documents to reconcile. The disambiguation has already happened.
Every field carries its origin and its date. An answer your model produces can be traced back to a page, which is the standard an IC or an LP will hold it to.
A model that can see how a value moved can answer questions a snapshot cannot. The time dimension is in the data rather than inferred from it.
REST, an MCP server for Claude, ChatGPT or your own agents, webhooks on every run, and Excel in your template when the output has to land in a workbook.
Kruncher exposes the whole layer through a Model Context Protocol server, so an agent can query companies, people, funds and the graph conversationally, inside your own governance boundary. Several customers run this way and never open the platform at all.
Read the MCP documentationThe thing customers say back to us most often is that this is a component they can delegate. Not a platform to adopt, not a workflow to move into, just a block that arrives maintained so the team can work on the part only they can do.
Document ingestion alone. Company analysis alone. The graph alone. Each is an API you call and a component Kruncher maintains on its own lifecycle, and everything you have already built can stay exactly where it is.
The valuation model, the pricing view and the investment decision stay with you, and Kruncher does not want visibility into them. That boundary is what makes the layer safe to adopt, and we state it before anyone asks.
When a reviewer picks a different value, that choice becomes the rule for the next document from that source, usually within the same week. The layer gets more accurate the longer you run it.
Deploy into your own AWS, Azure or GCP and nothing leaves your network. Run it against models your risk function has already approved, including open source and self hosted ones, with no egress to a vendor endpoint.
It usually works, and then it breaks. These are the four things that decide it, and none of them are about the prototype.
A platform you build once and run for five years is a different thing from this. Models, prompts and the formats your sources publish in move every few weeks, and a tool does not get better by sitting on a shelf. What ships today is behind by next quarter unless somebody is actively pushing it forward.
The first budget is rarely the last. The same line item returns every year and buys no new capability, only the right to keep what you already had. Teams who have been through it describe the internal roadmap as a standing argument about which fields to cut.
The engineers who enjoy building the first version rarely enjoy maintaining the fifth. Once the interesting part is done they go looking for the next interesting part, and that is the quiet moment an internal tool stops improving while still appearing to work.
Your model, your judgement and the decisions your firm takes are core, and nobody should build those for you. Entity resolution, twenty source integrations and a document pipeline that survives a live process are not core. They are the part worth building on top of.
The usual blocker, that your data ends up somewhere you cannot see, is a deployment question rather than a trade off. Run the whole layer inside your own cloud, against models your risk function has already approved. How that works.
Teams who never open the platform and run entirely on the API and the MCP server, with Kruncher as the substrate beneath their own agents.
Holdings re analyzed on a schedule so a fair value mark rests on sourced, timestamped evidence rather than the last reported figure.
Companies and the people behind them validated against auditable records, as a capability inside somebody else's product.
Every company and contact record kept current without anyone entering data, across a book too large to maintain by hand.
Bring the use case and we will tell you on the call which layers it needs, what stays yours, and how it deploys inside your boundary.