Guides / Pricing

How Much Does Custom Lead Generation Software Cost?

What actually drives the number: sources, formats, volume, verification depth and where the output has to land. Plus what a $1500 build really includes.

Published 2026-07-26 · 10 min read · by Lead Mining Company

Builds here start at $1,500. That number answers almost nothing on its own, because "lead generation software" describes a single scraper that pulls one county's new business filings and it also describes an engine that merges eleven sources, resolves duplicates across them, verifies contact data, and writes clean records into a CRM every morning. Those are not the same job. This guide explains what actually moves the price, what the entry build includes, and how to tell whether you should be paying for software at all.

What you are actually buying when you commission lead software

A lead system has four parts. Extraction gets raw records out of a source. Normalization turns inconsistent raw records into a consistent schema. Verification decides which records are real and current. Delivery puts them where you work.

Most buyers think they are paying for extraction. Extraction is usually the cheapest part. Pulling a table off a page or parsing a fixed-width state export is a solved problem, and an experienced developer does it quickly. The cost lives in everything after extraction: deciding that "Smith Roofing LLC" and "Smith Roofing, L.L.C." are one company, deciding that a phone number scraped in March is still connected in July, deciding what happens when a county changes its export format without warning.

When you compare quotes, compare on those four parts. A quote that only prices extraction is not cheaper. It is incomplete.

The six cost drivers that actually determine your quote

1. Number of sources

Cost does not scale linearly with sources. Two sources cost more than twice one source, because two sources introduce the reconciliation problem: the same business may appear in both, described differently. Ten sources are not five times two sources. They are a data-integration project.

2. Source hostility and format quality

Sources fall on a rough ladder. A published CSV or a documented API is the floor. A clean HTML table is close behind. Then paginated search forms that require a query per record. Then sites with session tokens, rate limits, or aggressive bot detection. Then scanned PDFs that need OCR, where accuracy work can exceed the entire rest of the build.

Ask where your sources sit before you budget. A prospect who says "it is all on the county website" often means a search form that only returns results one parcel at a time. That is a real build. See web scraping services for how those cases get handled.

3. Volume

Volume matters less than people expect up to a point, then matters a lot. A few thousand records a month runs on modest infrastructure with simple scheduling. Millions of records means queueing, proxy rotation, incremental sync instead of full re-pulls, and storage that behaves under load. The jump is not gradual. It arrives as an architecture change.

4. Verification depth

You can accept records as-found. You can validate format. You can validate deliverability and line status. You can validate against a second independent source. Each step adds engineering and, often, per-record vendor cost. Decide what a bad record costs you before deciding how much verification to buy. If your sales team burns ten minutes on every dead number, verification is the highest-return line in the quote. If you are running a low-touch email sequence, lighter verification may be rational.

5. Delivery target

A CSV in a folder is nearly free. A Google Sheet is close. A real CRM push means authenticating against their API, mapping your schema to their objects, handling their rate limits, deciding what happens on duplicate detection inside the CRM, and handling partial failures without silently dropping records. A hosted searchable database with filters and user accounts is its own application. Read custom lead database for what that involves.

6. Hosting and run cadence

A monthly run on a small server is cheap. Continuous monitoring with alerting on failure, retries, and a dashboard is not. Cadence drives infrastructure, and infrastructure drives your recurring bill.

The single fastest way to lower a quote is to cut sources, not features. Three well-chosen sources verified properly beat eleven sources dumped into a spreadsheet.

What $1,500 buys, and what it does not

The entry build is deliberately narrow so it can be delivered well:

  • One source, mapped and confirmed accessible before work starts.
  • A defined field list agreed in writing. Name, entity, address, phone, email, filing date, whatever the source actually carries.
  • Extraction plus normalization into that schema.
  • Baseline verification appropriate to the fields, so you are not handed obvious garbage.
  • One delivery target. A file drop, a sheet, or a single CRM endpoint.
  • A scheduled run and code you own.

What it does not include, honestly: multi-source merging, cross-source deduplication, third-party enrichment, a user-facing interface, per-user accounts, or an open-ended maintenance commitment. Those are real work and they are quoted as real work.

Most single-source jobs land at or near the entry price. Multi-source engines land materially higher, and hosted database products with logins and search land higher still. If a stated budget cannot cover the described scope, you will hear that before anyone starts, not after.

Why multi-source engines cost more than the sum of their scrapers

Say you want contractors who just pulled permits, cross-referenced with active business registrations, cross-referenced with property ownership. Three scrapers. Each is straightforward. The build is not, and here is the honest reason.

The permit lists a contractor name as typed by a clerk. The registration lists a legal entity name. The property record lists an owner who may be an LLC that owns nothing else. Matching these requires rules you have to invent for your specific data: how much name variance to tolerate, whether address is authoritative, what to do when two records agree on phone but disagree on name, which source wins on conflict.

Extraction is a known problem with known answers. Entity resolution is a judgment problem, and judgment has to be encoded, tested, and tuned against your actual records.

That tuning is where the hours go. It is also where the value is, because a merged, deduplicated, verified record is worth many times a raw scraped row. Scraper development is the input. The merge logic is the product.

Fixed price versus hourly, and why we quote fixed

Hourly protects the vendor. Fixed price protects you, but only if the scope is known. The problem with fixed-price software quotes is that nobody can price what they have not examined, so vendors either pad heavily or discover mid-build that a source is worse than advertised and come back for more money.

The way around it is a source-mapping phase first. Before quoting, each source gets opened and checked: what fields exist, how the pagination works, what defenses are present, what the record volume actually is, whether historical data is reachable or only current. That takes a short amount of time and removes most of the uncertainty. After mapping, a fixed price is honest, because the unknowns have been converted into known work.

If a vendor quotes a fixed price for a multi-source build without looking at the sources, be careful. Either they padded it or they will renegotiate later.

Ongoing costs nobody mentions in the sales conversation

The build price is not the total price. Budget for three recurring items.

Hosting. A small scheduled scraper runs on infrastructure costing very little per month. A hosted database serving users, storing millions of records, and running continuous jobs costs meaningfully more. Ask for an estimated monthly figure at quote time and ask what happens to it if volume triples.

Maintenance. Sources change. A county redesigns its portal, a state moves to a new vendor platform, an API version is retired. This is not a defect in the software. It is the nature of the work. You need either a maintenance arrangement or the willingness to pay for repairs as they arise. What you should not accept is a vendor who disappears and leaves you with code you cannot service. Owning the source code matters for exactly this reason.

Source access fees. Some counties and states charge for bulk data, per-record exports, or subscription access to their systems. Some charge nothing for browsing but charge for downloads. Some require an account and a signed data-use agreement. These fees are yours, not the developer's, and they should be identified during source mapping so they do not surprise you.

Build versus subscription: how to reason about it

Do not run this comparison on price alone. Run it on four questions.

Does a subscription tool actually cover your segment? The large lead databases cover common firmographic slices well. They cover event-driven niches poorly, because their business depends on breadth, not on your county's daily foreclosure docket. If an existing tool covers your segment, subscribe. That is cheaper and faster than anything custom.

What is the cost shape over time? A subscription is a recurring cost that continues as long as you need leads and typically rises with seat count and usage. A build is a larger cost once, plus a smaller recurring cost for hosting and maintenance. Somewhere on the timeline the lines cross. Where they cross depends entirely on your subscription price and your build price, so do that arithmetic with your real numbers rather than a rule of thumb. Run it over twenty-four months, not twelve, because most of the advantage of owning shows up in year two.

Who else has these leads? Everyone paying for the same subscription is calling the same list. A source-specific build gives you records at the moment they appear, from a source your competitors have not bothered to parse. That difference does not show up in a cost comparison and it is often the whole reason to build.

What happens if you stop paying? Subscription access ends and the data goes with it. Owned software keeps running and the accumulated database stays yours. That is a real asset with a real value, and it belongs in the comparison.

When you should not build custom software

Saying this costs us work, and it still needs saying.

  • You need a list once. If this is a one-time pull of a few thousand records with no ongoing need, buy the data or pay someone to do a manual extract. Software you run once is software you overpaid for.
  • The source does not exist. Some data people assume is public is not published, is not machine-readable, or is only available in person. No amount of engineering creates a source. If mapping shows the data is not reachable, the correct answer is no build, and you should hear that before you pay.
  • An off-the-shelf tool already covers your segment well. If a modest monthly subscription gives you the same records with the same freshness, subscribe. Come back if and when you outgrow it.
  • Your sales process cannot absorb more leads. If nobody is working the leads you have, more leads change nothing. Fix the follow-up first. Volume amplifies a working process and exposes a broken one.
  • The legal position is unclear. Some sources carry terms or licensing conditions that make automated collection a bad idea. Sort that out before commissioning, not after.

What to ask any vendor before you pay

Take this list to whoever you are considering, including us.

  • Did you open my sources before quoting? Ask what they found. A vendor who examined the source can describe its pagination, its fields, and its quirks. A vendor who did not will speak in generalities.
  • What exact fields will each record contain? Get the list in writing. "Contact information" is not a field list.
  • How is verification done, and what pass rate should I expect? They should be able to describe the method. Be suspicious of anyone promising a specific accuracy number before seeing your data.
  • Do I own the source code? If the answer is no, you are renting with extra steps, and you cannot hire anyone else to maintain it.
  • What is the estimated monthly hosting cost, and what raises it?
  • What happens when a source changes format? Ask specifically who pays and how fast they respond.
  • Are there source access fees, and are they in the quote?
  • What is out of scope? The most useful sentence in any proposal is the one listing what is not included.
  • Can I see it run before final payment? A live run on real records beats any demo.

The cheapest way to find out what your build costs

Send the sources. Not a description of the sources, the actual URLs or the name of the county system. Then say what fields you need on each record and where the records should end up. That is enough to map the sources and come back with either a fixed price or a straight answer that the build does not make sense for you.

If you want to scope your own project first, start with the field list. Write down every field you need and mark each one as required or nice-to-have. Most quotes shrink noticeably once the nice-to-haves are separated out, because half the cost in a typical build lives in three or four fields that turn out to be optional. Bring that list to a scoping conversation, or read how the process works on the custom lead generation software page.

More guides