Kruncher Logo
Build on Kruncher · The data layer

Build what makes you different. Leave the rest to us.

Ingestion, extraction, a resolved data model, a time series and a knowledge graph over millions of private companies. Reachable as an API and an MCP server, deployable inside your own cloud, and stopping exactly where your own judgement begins.

Millions
companies in the graph
~1,000
data points per company
20+
premium sources reconciled
14 steps
in the ingestion pipeline
The architecture

One resolved record, and everything else built on top of it

Your sources go in at the bottom. The ingestion layer and the time series engine turn them into one record per company, person and fund. The knowledge graph connects them. Workflows, analytics and your own agents sit on top.

The same diagram every Kruncher customer sees inside the product. The difference for a builder is that each layer is reachable on its own, so you can take the part you need and keep everything you have already built.

How the layer is built

Five things have to be true before data is worth reasoning on

Each one is a hard problem on its own, and each one is where an internal build tends to stall.

01
Ingest

Everything a company leaves behind, in one place

A fourteen step ingestion pipeline takes documents, email, calendars, call transcripts, your CRM, your data rooms and internal registers, alongside more than twenty premium databases and public registries across several languages. Nothing is read until the engine knows which document it is looking at and which entity the content belongs to.

  • Documents, decks, filings, data rooms and spreadsheets
  • Email, calendar, transcripts, messaging and address books
  • 20 or more premium sources and public registries, multilingual
02
Extract

Structure first, then figures

The structure is resolved before a single value is read, so the engine knows whether a number belongs to a vehicle, an asset, a period or a person. Every candidate value keeps the file, the page and the position it came from, and competing candidates are kept beside the preferred one with the reason it was preferred.

  • Entity and structure resolved before extraction begins
  • Source file, page and coordinates carried on every value
  • Competing candidates kept, never silently discarded
03
Model

One record, four kinds of truth

Roughly a thousand normalized data points per company, and each one labelled by where it came from: what the company says about itself, what the public record shows, what Kruncher estimates, and what your own team knows. Contradictions are resolved by source precedence rather than averaged, and a gap is flagged rather than filled.

  • What the company says, what is public, what we estimate, what you know
  • The latest and most authoritative source governs a contradiction
  • Missing stays missing, and says so on the page
04
Time series

Not a snapshot, a history

Every value is dated to the period it describes rather than the day it was read, so a March deck reporting December figures produces December facts. The record is a time series, which is what makes it possible to ask what changed between any two dates and to see the trajectory rather than the last reading.

  • Values dated to their period, not to the file
  • Full change history with the source behind each change
  • Diffs between snapshots drive 600 configurable signals
05
Knowledge graph

What surrounds the company, not just the company

Entity resolution collapses the same company, person or fund appearing under different names across every source into one node. Around it sit typed, directional edges: investor, partner, vendor, customer, competitor, founder, executive. Millions of companies, and every edge carrying the same provenance as any other data point.

  • One node per company, person and fund, across every source
  • Typed directional edges, each resolving to a source
  • Second order events become detectable rather than invisible
AI ready data

A model can only be as good as what it is reasoning on

AI ready means more than a clean JSON response. It means the disambiguation, the provenance and the time dimension are already in the data, rather than something your agent has to work out.

Resolved, not retrieved

An agent asking about a company gets one record, not twelve documents to reconcile. The disambiguation has already happened.

Sourced, so it can be checked

Every field carries its origin and its date. An answer your model produces can be traced back to a page, which is the standard an IC or an LP will hold it to.

Dated, so it can be reasoned over

A model that can see how a value moved can answer questions a snapshot cannot. The time dimension is in the data rather than inferred from it.

Reachable where your agents are

REST, an MCP server for Claude, ChatGPT or your own agents, webhooks on every run, and Excel in your template when the output has to land in a workbook.

The MCP server

Point Claude, ChatGPT or your own agents at it

Kruncher exposes the whole layer through a Model Context Protocol server, so an agent can query companies, people, funds and the graph conversationally, inside your own governance boundary. Several customers run this way and never open the platform at all.

Read the MCP documentation
Claude, ChatGPT and custom agentsTools and resources, not a scraped UIRuns inside your governance boundaryThe same data the platform reads
Focus on what matters

A building block, so your team spends its time where the value is

The thing customers say back to us most often is that this is a component they can delegate. Not a platform to adopt, not a workflow to move into, just a block that arrives maintained so the team can work on the part only they can do.

Take one component, not a platform

Document ingestion alone. Company analysis alone. The graph alone. Each is an API you call and a component Kruncher maintains on its own lifecycle, and everything you have already built can stay exactly where it is.

We stop where your edge starts

The valuation model, the pricing view and the investment decision stay with you, and Kruncher does not want visibility into them. That boundary is what makes the layer safe to adopt, and we state it before anyone asks.

Your corrections become the rules

When a reviewer picks a different value, that choice becomes the rule for the next document from that source, usually within the same week. The layer gets more accurate the longer you run it.

It stays inside your boundary

Deploy into your own AWS, Azure or GCP and nothing leaves your network. Run it against models your risk function has already approved, including open source and self hosted ones, with no egress to a vendor endpoint.

Buy versus build

Most teams we meet have already built a version of this

It usually works, and then it breaks. These are the four things that decide it, and none of them are about the prototype.

01

This is not a software build

A platform you build once and run for five years is a different thing from this. Models, prompts and the formats your sources publish in move every few weeks, and a tool does not get better by sitting on a shelf. What ships today is behind by next quarter unless somebody is actively pushing it forward.

02

The run costs more than the build

The first budget is rarely the last. The same line item returns every year and buys no new capability, only the right to keep what you already had. Teams who have been through it describe the internal roadmap as a standing argument about which fields to cut.

03

The people who build it move on

The engineers who enjoy building the first version rarely enjoy maintaining the fifth. Once the interesting part is done they go looking for the next interesting part, and that is the quiet moment an internal tool stops improving while still appearing to work.

04

Build what is actually core

Your model, your judgement and the decisions your firm takes are core, and nobody should build those for you. Entity resolution, twenty source integrations and a document pipeline that survives a live process are not core. They are the part worth building on top of.

 Build it yourselfBuild on Kruncher
Shipping a changeA feedback cycle measured in weeksThe same week
Keeping up with modelsYour team, every few weeksIncluded
Adding a fieldCost and run time blow outA configuration change
New source formatA rewrite of the parserA correction that becomes a rule
When the builder moves onThe tool stops improvingNothing changes
Cost after year oneThe same budget, every yearOne line item
What your team ownsThe maintenanceThe decisions

The usual blocker, that your data ends up somewhere you cannot see, is a deployment question rather than a trade off. Run the whole layer inside your own cloud, against models your risk function has already approved. How that works.

In the field

What teams have built on it

AI research stacks

Teams who never open the platform and run entirely on the API and the MCP server, with Kruncher as the substrate beneath their own agents.

Valuation and mark review

Holdings re analyzed on a schedule so a fair value mark rests on sourced, timestamped evidence rather than the last reported figure.

KYC, KYB and compliance

Companies and the people behind them validated against auditable records, as a capability inside somebody else's product.

CRM enrichment at scale

Every company and contact record kept current without anyone entering data, across a book too large to maintain by hand.

Secondary and fund of funds assessmentCorporate and competitive intelligenceDeal origination toolingLP portfolio transparency

Tell us which block you want to stop maintaining.

Bring the use case and we will tell you on the call which layers it needs, what stays yours, and how it deploys inside your boundary.