Open up a handful of Stocklake's signals and, until recently, you'd see the same three numbers on all of them — conviction, confidence, flag_score, each 0-10, each set by whichever AI screener wrote the idea. They looked identical across every source. They didn't mean the same thing on any two of them. An 8 from one screener might be a genuinely rare, well-earned call. An 8 from another might be what that screener says almost every single time — the AI equivalent of a restaurant handing out five-star reviews to everyone who orders tap water. We spent the last few weeks fixing that, and the result is signal_score: one number, computed the same way regardless of source, now live on the signals feed, on every news article, and on the deep-dive research tool. Here's what it actually is, and why it took more than "just average the three old numbers."
Stocklake runs more than a dozen independent screeners — some read price and volume patterns, some read news, some read insider filings. Each one is its own AI persona with its own voice, and each one hands back a conviction score the same way a person would: by feel, in round numbers. That's the part that broke down at scale. Ask any model for a 0-10 score often enough and it doesn't spread evenly across the scale — it clusters. On one of our busiest sources, out of nearly 1,600 recent calls, 74% used the exact same conviction score. Not "mostly high," not "mostly low" — the identical number, three out of every four times.
That clustering makes a single fixed cutoff — "only act on anything scoring 8 or above" — nearly meaningless. If three-quarters of a source's calls already sit at 8, the cutoff isn't filtering anything; it's just describing what that source always says, the same way a doctor who diagnoses every patient with "mild concern" isn't triaging, he's just talking. We needed a way to tell a routine 8 apart from a genuinely rare one, and to do it the same way for every source, not with a different hand-tuned rule per screener.
signal_score is a single 0-100 figure built from three separate reads on the same idea, blended together:
1. The AI's own call. The original conviction/confidence/flag_score axes are still collected — they're real input, just not the final word anymore.
2. How unusual this call is, for this source. Instead of comparing an 8 from one screener to an 8 from another, we compare it to that same screener's own history. A conviction=8 from a source that almost always says 8 barely moves the needle. A conviction=8 from a source that's said it twice this quarter is a real outlier, and scores accordingly.
3. One real, checked fact. Every screener is asked to name a specific, concrete factor behind its call — a technical reading, a fundamental detail — and we look that value up against live data ourselves rather than trusting whatever number the AI typed. If it names something we can't verify, that piece of the score is simply dropped, not guessed at. This is the piece that actually breaks the round-number habit, because it's anchored to a real, continuously-varying number instead of a model's gut feel.
A 10-point scale only has 10 places to land, and a handful of round numbers (6, 7, 8) already absorb most of a model's actual output — it's a parking lot with three spots doing the work of thirty. A 0-100 scale gives the same underlying judgment room to actually spread out, which matters most exactly where it counts — separating a good call from a great one, instead of both landing on the same "8."
The old world used one flat cutoff for every screener, regardless of how good that screener's calls had actually turned out to be — a rookie and a decade-long veteran, held to the identical bar, on day one. That's backwards: a source with a real, proven track record should get more benefit of the doubt than one that's never been checked. Now each source earns its own bar, set from its own real results and re-checked on a schedule — not hand-picked once and left alone. A source with a strong history can clear the bar at a more modest score; a source with no track record yet, or a poor one, needs a much stronger score to earn a look. Nothing about this is manually curated per source going forward — the bar moves as the evidence does.
| Before | Now | |
|---|---|---|
| Scale | Three separate 0–10 numbers | One 0–100 number |
| Set by | The AI, from scratch, every call | The AI's call + its own track record + one verified fact |
| Comparable across sources? | No | Yes — same method everywhere |
| Bar to clear | One fixed number for everyone | Each source earns its own, from its own results |
We built signal_score to decide, internally, which ideas were worth surfacing at all — that part had already been running for a while before any of it was public. It's now a field on the signals feed itself, on every news article we score, and inside the deep-dive research tool — so wherever you're pulling an idea from, you get this number alongside it, from every source, on the same scale. Practically: instead of learning the personality of a dozen-plus different screeners to figure out what their numbers actually mean, you get one already-normalized figure you can sort, filter, and compare on directly — no source-specific calibration required on your end.
ai_score, and insider & institutional activity is on the same signal_score everything else uses. Same 0-100 scale, same four bands, everywhere it shows up. The full story on why two scores, not one →signal_score answers "how much real, checkable evidence backs this specific call, relative to what this source usually says" — not "will this idea play out." Those are related questions, not the same one. It's still fundamentally shaped by an AI's own read of the situation, the per-source track record is only as good as the history behind it (thin for newer sources), and the whole thing is deliberately built to keep re-checking itself against real results rather than being tuned once and trusted forever. Treat a high score as "unusually well-evidenced," not as a guarantee.