We sell done-for-you lead generation. Then we had the same problem ourselves — a list of 374 companies and no way to contact anyone on it. So we pointed 50 AI agents at it and left them running. This is the full log.
In July we filtered 1,463 recently-funded YC companies down to 374 that matched our buyer. Then we stalled.
The list held companies, not people. No names. No emails. No profiles. Paid enrichment tools quoted real money for the rest.
By hand, we qualified ten of them in three weeks. At that rate the list would take two years.
Not one long job — nine multi-agent workflows, two power failures, four scraper rewrites.
Handed a 200MB transcript of a dead session and told to catch up. It parsed 40,603 lines and reconstructed a month of context — who we had spoken to, what was promised, what was actually delivered.
Our internal tools still ran inside a dead brand's infrastructure. Forked, database migrated, domain moved, single sign-on wired. 153 links, 96 videos, 6 leads carried across. No downtime.
One of the agents found a route to the people behind the companies that the rest of the market has walked past. Six minutes later it had the whole list.
730 founders. 725 LinkedIn profiles. 99.7% coverage.
One scraper survived because it saved its progress. The other didn't — it wiped its own table on every restart. Four failed runs before the real bug surfaced: 44 minutes of CPU burned by pattern-matching against minified JavaScript. Never a network problem at all.
Then two more were told to attack the analysis. They proved the first read wrong on four counts — including a claim our investor had killed a plan he was never actually told about.
Each round attacked the same 374 companies with a technique the previous rounds hadn't tried. Ran unattended for ten hours.
Rounds 1–8 replayed from cache instantly. Only unfinished work re-ran. Cost of the second crash: nothing.
Ten methods, run in sequence, each one told what the previous nine had already tried. The top three carried 63% of everything found. What they are is the part we keep.
Rounds 6 and 7 were duds — and that matters. The run was originally designed to stop after two weak rounds in a row. That logic was wrong: each round is a different door, so a locked one tells you nothing about the next. We removed it. Rounds 8 and 10 went on to find 295 more contacts.
Every corner of a company's web presence, not just the front door.
One agent, one script, fifteen thousand requests, four minutes.
What a company's site used to say is often more useful than what it says now.
Logged out, no scraping tool — enough to rank every content format by what actually performs.
Where the compute actually went was the opposite of what I expected.
An agent that writes a script and reads back a summary is doing arithmetic. The heavy lifting happens outside the model, and it scales almost without limit.
The expensive hour was six agents reading one 62-minute transcript, twice over, then arguing about it. No script compresses that — every word has to be read and weighed.
The machine can do a hundred thousand things. Deciding which hundred thousand is the job.Which is why judgement, not volume, is where the money goes
Published because the wins mean nothing without these.
Reported as fact. An adversarial agent then searched both transcripts: he was never told the plan. He'd rejected one tactic, and his silence on the rest was read as a verdict.
It was. One folder over from where the search ran.
Stated as fact, never tested. It was wrong — a step that failed under one model went through under another.
One missing line of code rendered it as a shrunken desktop page. We read everything on our phones. That single omission explained months of preferring PDFs to our own tools.
Startups copy each other's boilerplate word for word, so a page can hand you somebody else's contact details. One yielded an address for the California Department of Consumer Affairs. Another yielded ten law firms. 25 rows quarantined rather than shipped — this is exactly the failure a cheap list ships to you and never mentions.
Compiling everything into one outreach file was refused — the shape resembles mass data-harvesting. Legitimate business, blunt rule. We ran the last command ourselves rather than route around it.
Reproducible the day the next YC batch lands.
The most valuable output of 43 hours wasn't a contact. It was an agent proving the analysis wrong.
Every significant finding got a second agent whose only job was to refute it. That pass caught four real errors before they reached a decision — including one that would have walked me into a VC meeting with the wrong story.
A quiet round tells you nothing about the next one. A locked door tells you nothing about the door beside it.Why we kept going after two failed rounds — and found 295 more contacts
Nothing here is about my industry. It is all about not losing two days of machine work.
Written up by Faris Irfan, who mostly just kept asking why it had stopped.