SendVC

Research · measured 2026-10-09

State of Early-Stage VC 2026: how investors actually describe themselves

Across 5,000+ active investor records, investors describe what they fund using 2,500+ distinct labels — and those labels are so inconsistent that counting the single most common spelling of a category undercounts the real number of investors in it by as much as 43.2%. Anyone quoting you a clean figure for “how many investors fund X” is almost certainly reading one spelling of it.

What this is measured over

Every figure below is computed at build time from SendVC's active investor table. Nothing is estimated, extrapolated or typed in by hand. The table is early-stage-weighted by construction, so read it as a description of the early-stage market rather than of venture capital as a whole.

5,000+
Active investor contacts
94.9%
Have focus areas recorded
93.7%
Have a location recorded
43.2%
Have portfolio data recorded
8.7%
Have a cheque size recorded

The last figure is the one to keep in mind. Cheque size is recorded for 8.7% of investors, so this report publishes no median cheque size and no cheque-size distribution anywhere. A median computed from coverage that thin would look measured without being measured. Weak denominators are printed here rather than hidden, because a report that only shows its strong numbers is not a report.

Finding 1: the vocabulary problem is bigger than the data problem

2,500+ distinct focus labels appear across the table, of which 1,600+ are used by exactly one investor. The twenty most common labels account for only 43.5% of all label usage. In other words, there is no shared taxonomy: most investors are describing themselves in language nobody else uses in quite the same way.

This is why keyword-matching a deck against an investor list performs so poorly, and why the fix has to be semantic rather than lexical. Two investors who fund exactly the same thing routinely have zero words in common in how they say it.

The inconsistency reaches the firm names themselves: a fraction of a percent of firms are stored under more than one capitalisation of their own name, so the distinct-string count is slightly higher than the true number of firms. That delta is tiny, and it is mentioned precisely because it is tiny — it is the cleanest available demonstration that even the least ambiguous field in an investor database does not deduplicate itself.

Most common focus labels

Finding 2: stage labels undercount themselves

The fragmentation is measurable most precisely on investment stage, because stage has a small, well-understood set of real values that the raw strings should map onto — and mostly do not. For each canonical stage below, “spellings” is the number of distinct raw strings that fold into it, and “undercount” is how far the single most common spelling falls short of the folded total.

StageInvestorsSpellingsMost common spellingUndercount
Early-stage600+82Early-stage43.2%
Seed450+68Seed37.7%
Growth250+70Growth67.6%
Pre-seed150+27Pre-seed48.7%
Series A100+20Series A44.1%
Angel40+15Angel Network68.1%
Series B+under 258Series B66.7%
Late-stageunder 2512Late-stage68.8%

Read the last column as: if you queried this database for the most common spelling of a stage and reported that number, this is the share of matching investors you would have missed.

Finding 3: only a minority of investors state a stage at all

1,100+ investors (20.5%) state an investment stage in any form. The rest describe sectors and say nothing about when they enter. That is the single biggest reason stage-based investor lists are unreliable: for four investors in five, the stage is simply unknown rather than excluded.

Among those who do state one, 500+ state more than one — “seed to Series A” is a range, not a point. Stage shares therefore add to more than 100%, and the table below is deliberately not normalised to hide that.

StageInvestorsShare of those who state a stage
Early-stage600+53.8%
Seed450+40.9%
Growth250+26%
Pre-seed150+13.5%
Series A100+12.7%
Angel40+4.2%
Series B+under 251.9%
Late-stageunder 251.4%

Finding 4: location is recorded far more often than it is usable

93.7% of investors have a location on file, but those values take 1,500+ distinct forms — the same city written as a bare name, with a state, with a country, or with the parts reversed. After folding the recognisable variants together, only 2,100+ records (38.6%) resolve to a canonical city.

The gap between those two percentages is the point. A geographic investor list built by grouping on the raw column would split one city across several thin buckets and silently drop the rest.

Method, and what it cannot tell you

Figures are computed directly from SendVC's active investor records on the date shown at the top of this page, and regenerate whenever the page rebuilds. Canonical stages and cities are produced by the same folding functions the rest of the product uses, so the numbers here match what the product acts on rather than describing a separate research dataset.

Three limits worth stating plainly. The table is early-stage-weighted, so it is not a census of venture capital. Coverage varies by field, and every percentage above is reported against the field's own denominator rather than a flattering one. And no claim is made here about outcomes — this report describes how investors present themselves, not who funds what or how often anyone replies.

The practical version of this

If the vocabulary is this fragmented, building an investor list by filtering on labels will miss most of the right investors. SendVC matches a pitch deck against5,000+ verified contacts semantically instead — plus additional investors it finds and validates across the web for that specific startup — then writes a personalized email for each match, which you approve before anything sends.

Cite as: SendVC, “State of Early-Stage VC 2026: how investors actually describe themselves”, https://www.sendvc.app/reports/state-of-early-stage-vc-2026, measured 2026-10-09.