Per-publisher metric breakdown

This report makes the cross-publisher comparison explicit: the same evaluation contract should expose different publisher shapes without pretending they are the same newsroom. ADR-0006 calls this the analytical contract. Analysts should be able to query a common schema over EB-NeRD, Adressa, and MIND, but the common schema should not wash away the local behaviour of each source. The heatmap below is therefore a comparison surface, not a leaderboard.

ADR-0009 sets up the comparison by putting click metrics beside editorial metrics. That is why the rows here include NDCG@10, MRR, hit rate@10, diversity, coverage, and sensitive exposure. A normalized benchmark would be easier to sort. It would also be less useful, because it would hide the reason an editor cares about cross-publisher checks in the first place: the same policy can create different tradeoffs depending on the source material and audience pattern.

EB-NeRD's curve is steepest because it begins with the most room to move on the editorial side of the evaluation. In the fixture, its click metrics are lower and its baseline diversity and coverage are lower, so a diversity-forward configuration reads as a large editorial shift rather than a small polish. That does not make EB-NeRD worse. It makes the tradeoff more visible. The analyst question is whether that steeper movement is acceptable for the product surface, not whether the source should be forced to look like another dataset.

Adressa's curve is flattest because the evaluated shape is already less compressed. Its averages sit in the middle: stronger click metrics than EB-NeRD, less ceiling pressure than MIND, and a diversity/coverage posture that does not require the same dramatic correction. A flat curve can be good news if it means policy changes do not create unstable metric swings. It can also hide risk if the organisation stops asking whether the flatness is caused by a small fixture, the source mix, or a constraint that is not strong enough to matter.

MIND behaves differently because it begins near the top of the click metrics, so the comparison reads less like rescue and more like governance. A high NDCG@10 value does not end the evaluation. It raises the standard for explaining what editorial value the platform is preserving while keeping click performance strong. The point of including MIND is not to crown it; the point is to show that even a favourable benchmark still needs the editorial metrics beside the click metrics.

The volume table keeps the comparison honest. In a full warehouse, impression counts, user counts, article counts, and time ranges would explain part of the metric shape. In this clean-room fixture, the counts are intentionally tiny, but the contract is the same one the deployed analyst surface reads. That is why the report keeps the table visible instead of hiding it behind a summary sentence. An analyst needs to know the evaluation shape and the population shape before making a claim about publisher behaviour.

That population context is also where a future production analyst would start looking for counter-explanations. A curve may be steep because the editorial policy changed, because the catalogue mix changed, because traffic concentrated around one story cluster, or because the fixture is too small. The shared contract makes those explanations testable.

The cross-publisher lesson is therefore cautious. Shared contracts make comparison possible; they do not make publisher context disposable. EB-NeRD shows the largest editorial movement, Adressa shows the gentlest movement, and MIND shows what governance looks like when click metrics are already strong. That is the kind of comparison an editorial platform should invite: specific, qualified, and grounded in the same tables the rest of the system uses.