Home/Case studies/Lead qualification

Case study 05 / B2B lead generation

5M+ companies processed.
98% accurate.

A production multi-agent pipeline qualifies companies across industries and overlapping niches while scaling model capability with the complexity of each case.

5M+ companies processed92% lower model costHuman review for rare exceptions
01

98%

classification accuracy in production

02

−92%

model cost through complexity routing

03

5M+

companies processed across industries

04

3 tiers

of progressively capable agents

The classification problem

Classify 5M+ companies.
Separate 40 niches.

A B2B outreach program needed to qualify tens of thousands of companies against dozens of narrow, sometimes overlapping sub-niches. Generic provider filters routinely pulled adjacent but incorrect businesses into the same list.

The issue was classification precision. One company could legitimately qualify for Segment A and C but not B. A shared classification field silently overwrote that reality and contaminated campaigns.

The hardest errors came from companies offering many different services. Their broad language looked relevant to several niches, which made disciplined context and confidence routing more important than simply choosing a larger model.

The system

Route by complexity.
Cut model cost 92%.

Every tier has a high-confidence exit. Most records resolve from the company description. Only harder cases earn website tools, web search, broader context, or human judgment.

Qualification run / confidence escalation

Every tier has a confident exit

5M+ processed98% accurate
ONE COMPANY × ONE SEGMENTCONFIDENT RESULT AT ANY TIERINPUTAcme Incsegment candidatePASS 1 / CHEAPESTStrict text classifierCompany description onlyNarrow context, fast verdictHIGH CONFIDENCELOW CONFIDENCETIER 2 / CAPABLEWebsite investigationSitemap + targeted pagesFailure type preservedHIGH CONFIDENCELOW CONFIDENCETIER 3 / MOST CAPABLEAlternate-source agentWebsite + web searchDirectories, news, LinkedInHIGH CONFIDENCELOW CONFIDENCERARE EXCEPTIONHuman reviewWRITE SEGMENT MATCHQualified or excludedVerdict + confidence + evidenceIndependent for every segment98% accurate in productionCOMPLEXITY CONTROLS CAPABILITYSmall context and cheap model first.Tools and broader context only on escalation.92% lower model cost

Data and context architecture

Isolate each company.
Preserve every verdict.

A junction table stores one independent company-to-segment verdict per row. Each execution receives one company, one segment definition, the relevant qualifying and disqualifying criteria, and only the evidence gathered for that case.

Multi-label by design

A company can qualify for A and C but not B without overwriting any result.

Context managed tightly

Every tier sees only the criteria, evidence, and tools required for its decision.

State carried forward

Confidence, sources, and the precise reason for escalation travel with the record.

Production engineering

Process millions.
Contain four failure modes.

01Cross-item contaminationAgents were isolated into strict one-company executions so one record could never leak into another.
02Runaway tool useExplicit stopping conditions and tool budgets taught agents when a blocked page was a valid conclusion.
03Silent field lossCriteria were preserved across every escalation so later agents always knew exactly what they were evaluating.
04Index misalignmentFiltered branches were keyed to actual records rather than fragile array positions.

Cost follows complexity

Start with cheap models.
Spend 92% less.

A flat architecture would send every record through the most capable agent with website tools, web search, and a large context window. This system starts with the smallest useful model and narrowest useful context. Capability expands only after a low-confidence result.

Pass 1

Text only

Description, segment definition, and strict output schema.

Escalation

Tools on demand

Website access and external search appear only when the evidence requires them.

Measured result

−92%

Model cost versus treating every company as a maximum-complexity case.

Human review became an exception path.

It was reserved mainly for scraper-blocked websites and genuinely unusual companies the evidence could not resolve. The automated system produced accurate results around 98 percent of the time.

The outcome

98% accurate.
92% less model cost.

Map a qualification system