The Middle Band: Designing a Review Queue Operators Actually Use
Scores between 40 and 59 are not a scoring failure, they are a routing decision. Treating manual review as a first-class state changed the qualification machine.
The first version of our qualification stage had two outcomes. Above the threshold a publisher was queued for submission; below it, rejected. It was clean, it was easy to explain, and it was throwing away a meaningful slice of usable inventory every single run.
The publishers being discarded were not bad. They were ambiguous — and ambiguity is information, not noise.
Why the band exists
The authority score compresses eight independent signals into one number, and compression loses the shape of the input. Two publishers can both score 52 for completely different reasons: one is a well-run niche directory that is simply young and thinly indexed, the other is an ageing general directory with strong history and a spam profile that is starting to turn. Averaged into a single integer they are indistinguishable. Looked at as vectors they are obviously different decisions.
Review is a state in the machine, not a pause in it. A publisher sitting in review does not block the campaign, does not hold a worker, and does not stop anything else from being submitted. The rest of the run proceeds around it, and an operator's verdict simply moves it onto one of the existing paths.
A queue people use has three properties
It is bounded, it is explained, and it is fast. Bounded means the band is narrow enough that the queue is a handful of items per campaign rather than a second inbox — if a fifth of discovery lands in review, the thresholds are wrong and the queue is just deferred rejection.
Explained means the item shows the signal breakdown, not the total. The operator sees that index status and technical health are strong while spam signals and category fit are weak, and that is a decision they can make in seconds. Handing them the number 52 and a link is handing them the entire original problem.
An operator will not review a queue that shows them a number. They will review one that shows them a disagreement.
Fast means the verdict is two clicks and an optional reason code, with the publisher's live page one click away. Anything that requires typing will not get done on a Friday afternoon, and a queue that is not cleared is worse than no queue at all — it is a growing pile of inventory nobody is deciding about.
Every verdict is calibration data
This is the part we did not anticipate. Each review stores the operator's decision and reason code against the full signal vector that produced the score. That is a labelled dataset generated by the ordinary operation of the product, on exactly the examples where the model is least certain — which is where labels are worth the most.
When we re-weight the score, the review queue is the regression set. The auto-qualified and auto-rejected tails tell us almost nothing, because everything agrees about them. The middle band is where the weights are actually decided, and it took building the queue for the wrong reason to notice that.
The queue must be allowed to be empty
One last design rule. As calibration improves, the band should shrink and the queue should thin out. A review queue that grows with every campaign is not a safety valve, it is a design defect wearing one. We track queue volume as a percentage of discovered publishers per campaign and treat a rising trend as a scoring bug rather than an operations problem.