LiveCodeBench accuracy vs. contest date

A data-contamination check for the Qwen3 thinking-rollouts models. Rows are difficulty (all / easy / medium / hard); columns are pass@1 / pass@10 / pass@50. Every model is a line (colour = size, think solid · nothink dotted, house style from thinking.html), with 95% bootstrap-CI bands over 100 samples/problem. The dropdown switches temperature. If a model memorized LiveCodeBench problems seen during pretraining, its pass@k should look inflated on earlier contests and fall on later ones it could not have seen. Per-cell problem counts are on hover (n); thin difficulty×quarter cells (≥15) have wide CIs.

No official cutoff to split on. Qwen3 publishes no pretraining/knowledge cutoff — verified in the Qwen3 Technical Report and the Qwen3-8B model card (an informal community thread guesses ~mid-2024). So this shows the accuracy trend across contest time — read a downward step (seen consistently at 2024-Q3, concentrated in medium/hard) as the likely contamination boundary.

Data coverage — LiveCodeBench problems by contest date

Bars are binned by quarter (matching the accuracy plot's points); the rug beneath each panel marks every individual problem's contest date. A sanity check on how much data sits behind each point above — note the thinner difficulty×quarter cells.