An AI radar that finds clients: the hard part is not finding, it is invented quotes
The system reads open boards and brings you enquiries. The dangerous failure is not missing one, it is bringing you words nobody wrote. Here is what actually fixes that.
Can you find clients automatically with AI?
You can, and the finding is the easy half. The hard part is stopping the model from paraphrasing someone's post: it will happily produce a plausible quote that was never on the page. The fix is not a better prompt but a check — every stored quote is matched verbatim against the fetched page, and one that does not match never reaches the operator.
The system reads open boards, finds people describing a problem right now, and brings them to you. It sounds like a scraping problem. In practice scraping is the dull part of it.
Where systems like this break
Picture the record: a name, the board, a link, and a quote — the words the person used to describe what they need. The quote is what you read in two seconds to decide whether to reply.
And the quote is what a model invents most readily. It read a long thread, understood the gist and produced a tight formulation: fluent, convincing, and absent from the page. Not because it malfunctioned, but because that is what it was asked to do — “extract the request”.
The consequences are unpleasant. You reply to someone quoting words they never wrote. At best that reads as careless; at worst like a message meant for somebody else.
What actually fixes it
Not the prompt. However firmly you write “quote verbatim”, only code can check it.
There is a step that takes every stored quote and looks for it in the text of the source page. Not something similar — the same string. If it is not found, the record never reaches the operator.
It is a cheap check with a large effect. It turns “the model usually quotes correctly” into “the record either holds the person’s real words or does not exist”. The difference between those two states is the difference between a tool you trust and a tool you re-check by hand, which is to say a useless one.
A funnel from cheap to expensive
The second thing that decides whether a system like this survives is money.
Send every page you find to a large model and the bill grows faster than the value. So the checks are stacked from cheap to expensive: a blunt reject on the URL and the title, then a fast keyword pass, then a semantic gate on embeddings, and only what survives all of that reaches the model.
The order matters more than the contents. Each stage throws away most of the flow for pennies, and the expensive check works on the remainder rather than on the internet.
The engine and the niche are different things
The third decision is not visible at first, but without it the system lives in one niche and dies with it.
There is no mention of any particular trade inside the engine. The queries, keywords, the phrases that mark a buyer, the list of boards, the thresholds and the extraction profile all live in a separate pack for one niche. The engine knows nothing about them.
So moving the radar to another trade is a new pack, not a new system. For me that is also the answer to whether this can be sold: the engine is the product, the pack is the setup.
What to take from this
Three rules that apply to any system where a model reads someone else’s text.
Verify anything the model quoted against the source, in code. Stack the checks so the expensive one works on what is left. And keep the knowledge about a subject in data, not in code.
The radar is running, and the case shows the records and how it is put together. If you want something similar for your own niche, get in touch.