Acceptance rate (Z) and accuracy ceiling (skyline) of two potentials — the
no-runtime-error potential (LCBRuntimeNoErrorPotential) and the public-test potential
(LCBPublicTestPotential) — under exact posterior sampling, across 12 models and 6 temperatures
(71 cells, ~4.2M rollouts).
Two potentials, two constraints. Each rollout y for question x defines a validity indicator φ(y); rejection sampling targets π(·|x) ∝ p0(·|x)·φ(·) — sample from the model, keep the valid ones.
Pooled over all rollouts, Z=(∑valid)/(∑N), skyline=(∑correct)/(∑valid)=pass@1/Z, so pass@1 = Z × skyline holds on the plotted numbers. Correct = the LCB evaluator’s verdict over all tests (public + private).
No-error potential (LCBRuntimeNoErrorPotential): φ = the code runs on the public inputs with no runtime error / timeout. A program that runs but answers wrongly is still valid.
Public-test potential (LCBPublicTestPotential): the strict constraint is “passes all public example tests”. Because the potential is soft (finite penalty per failed public test), every public-test section is indexed by k = the max number of failing public tests allowed: φk=1[n_failed≤k]. k=0 is the strict potential; raising k relaxes it, and at k=max#public-tests every rollout is accepted (Z=1, skyline=pass@1). A correct solution passes every public test, so correct ⇒ φk and pass@1 = Zk × skylinek at every k.
Why “skyline”? It is the accuracy of ideal constrained generation: if you could sample exactly from π (enforce the constraint and nothing else), this is how often the result would be correct — a reference for judging approximate inference methods (importance sampling, SMC, fine-tuning) that target the same π. Reported pooled over problems so pass@1 = Z × skyline holds; shaded bands are 95% bootstrap CIs (2000 resamples of the problem set).
LiveCodeBench, 12 models × 6 temperatures (71 cells, ~4.2M rollouts); both potentials run literally (forkserver, 6 s per-test timeout). Instances with no public tests excluded from the public-test analysis.
Part 1 — no-error potential
skyline against temperature
skyline = (correct)/(runs-without-error), pooled. Wrong-but-running answers dilute it.
skyline by model and temperature
model
t=0
t=0.2
t=0.4
t=0.6
t=0.8
t=1.0
Llama-3 8B
0.138
0.131
0.119
0.106
0.085
0.063
Llama-3 8B Instruct
0.202
0.199
0.195
0.189
0.183
0.180
Llama-3.1 8B
0.132
0.127
0.119
0.106
0.088
0.068
Llama-3.1 8B Instruct
0.229
0.224
0.222
0.216
0.206
0.203
Llama-3.2 1B
0.027
0.022
0.018
0.013
0.009
0.005
Llama-3.2 1B Instruct
0.098
0.095
0.088
–
0.067
0.054
Qwen2.5-Coder 3B
0.273
0.273
0.269
0.257
0.227
0.190
Qwen2.5-Coder 3B Instruct
0.297
0.300
0.298
0.289
0.282
0.273
Qwen2.5-Coder 7B
0.321
0.319
0.315
0.299
0.271
0.240
Qwen2.5-Coder 7B Instruct
0.399
0.394
0.390
0.387
0.381
0.377
DeepSeek-Coder 6.7B
0.185
0.187
0.182
0.173
0.159
0.139
DeepSeek-Coder 6.7B Instruct
0.242
0.234
0.227
0.227
0.224
0.220
Z, the acceptance rate (normalizing constant)
Z = pooled proportion of rollouts that run without error; falls toward t=1.0.
by difficulty, choose a model
No-error skyline and Z split by problem difficulty for a single model. Shaded bands are 95% bootstrap CIs over that difficulty’s problems.
skyline & Z by difficulty and temperature (selected model):
difficulty
metric
t=0
t=0.2
t=0.4
t=0.6
t=0.8
t=1.0
easy
skyline
0.340
0.323
0.294
0.260
0.199
0.136
Z
0.878
0.875
0.863
0.811
0.721
0.494
medium
skyline
0.028
0.029
0.022
0.018
0.012
0.007
Z
0.773
0.786
0.770
0.718
0.612
0.372
hard
skyline
0.000
0.000
0.001
0.000
0.001
0.000
Z
0.682
0.681
0.651
0.571
0.440
0.232
difficulty
metric
t=0
t=0.2
t=0.4
t=0.6
t=0.8
t=1.0
easy
skyline
0.477
0.483
0.474
0.460
0.441
0.417
Z
0.899
0.892
0.879
0.855
0.821
0.769
medium
skyline
0.056
0.049
0.049
0.047
0.044
0.041
Z
0.845
0.823
0.806
0.781
0.740
0.651
hard
skyline
0.015
0.005
0.002
0.002
0.002
0.003
Z
0.667
0.676
0.676
0.659
0.614
0.523
difficulty
metric
t=0
t=0.2
t=0.4
t=0.6
t=0.8
t=1.0
easy
skyline
0.330
0.328
0.303
0.264
0.211
0.148
Z
0.866
0.875
0.872
0.833
0.743
0.527
medium
skyline
0.023
0.014
0.014
0.013
0.009
0.005
Z
0.780
0.790
0.785
0.739
0.630
0.385
hard
skyline
0.000
0.000
0.000
0.000
0.000
0.000
Z
0.662
0.685
0.657
0.596
0.463
0.244
difficulty
metric
t=0
t=0.2
t=0.4
t=0.6
t=0.8
t=1.0
easy
skyline
0.509
0.517
0.519
0.501
0.468
0.424
Z
0.933
0.929
0.926
0.914
0.892
0.830
medium
skyline
0.084
0.076
0.069
0.066
0.055
0.046
Z
0.863
0.855
0.844
0.820
0.775
0.638
hard
skyline
0.022
0.016
0.014
0.011
0.011
0.010
Z
0.677
0.724
0.717
0.689
0.606
0.422
difficulty
metric
t=0
t=0.2
t=0.4
t=0.6
t=0.8
t=1.0
easy
skyline
0.078
0.060
0.050
0.033
0.024
0.012
Z
0.592
0.608
0.588
0.542
0.462
0.297
medium
skyline
0.000
0.000
0.000
0.001
0.000
0.000
Z
0.607
0.593
0.571
0.532
0.456
0.285
hard
skyline
0.000
0.000
0.000
0.000
0.000
0.000
Z
0.475
0.437
0.402
0.363
0.295
0.168
difficulty
metric
t=0
t=0.2
t=0.4
t=0.6
t=0.8
t=1.0
easy
skyline
0.217
0.216
0.203
–
0.149
0.112
Z
0.773
0.788
0.765
–
0.635
0.455
medium
skyline
0.017
0.018
0.015
–
0.009
0.005
Z
0.653
0.674
0.648
–
0.502
0.327
hard
skyline
0.012
0.007
0.006
–
0.004
0.003
Z
0.434
0.487
0.476
–
0.355
0.195
difficulty
metric
t=0
t=0.2
t=0.4
t=0.6
t=0.8
t=1.0
easy
skyline
0.565
0.565
0.553
0.525
0.459
0.367
Z
0.908
0.912
0.907
0.889
0.838
0.679
medium
skyline
0.146
0.145
0.141
0.134
0.110
0.072
Z
0.816
0.820
0.813
0.780
0.711
0.504
hard
skyline
0.015
0.016
0.016
0.013
0.012
0.006
Z
0.677
0.673
0.653
0.627
0.548
0.335
difficulty
metric
t=0
t=0.2
t=0.4
t=0.6
t=0.8
t=1.0
easy
skyline
0.569
0.592
0.592
0.578
0.560
0.530
Z
0.945
0.924
0.922
0.915
0.900
0.866
medium
skyline
0.181
0.176
0.172
0.162
0.153
0.134
Z
0.816
0.830
0.828
0.812
0.781
0.705
hard
skyline
0.008
0.015
0.015
0.013
0.012
0.009
Z
0.611
0.627
0.633
0.625
0.591
0.495
difficulty
metric
t=0
t=0.2
t=0.4
t=0.6
t=0.8
t=1.0
easy
skyline
0.648
0.644
0.627
0.591
0.528
0.427
Z
0.954
0.953
0.951
0.937
0.891
0.727
medium
skyline
0.202
0.203
0.202
0.186
0.159
0.124
Z
0.892
0.906
0.890
0.853
0.777
0.540
hard
skyline
0.014
0.022
0.024
0.023
0.018
0.014
Z
0.742
0.761
0.748
0.704
0.606
0.334
difficulty
metric
t=0
t=0.2
t=0.4
t=0.6
t=0.8
t=1.0
easy
skyline
0.758
0.729
0.719
0.711
0.700
0.689
Z
0.954
0.970
0.969
0.964
0.959
0.945
medium
skyline
0.293
0.299
0.298
0.297
0.288
0.274
Z
0.888
0.890
0.879
0.874
0.860
0.828
hard
skyline
0.033
0.035
0.032
0.030
0.026
0.026
Z
0.763
0.762
0.757
0.742
0.724
0.673
difficulty
metric
t=0
t=0.2
t=0.4
t=0.6
t=0.8
t=1.0
easy
skyline
0.412
0.414
0.400
0.385
0.352
0.296
Z
0.958
0.946
0.938
0.922
0.875
0.723
medium
skyline
0.086
0.095
0.093
0.081
0.065
0.049
Z
0.877
0.868
0.859
0.833
0.766
0.577
hard
skyline
0.006
0.007
0.006
0.004
0.004
0.002
Z
0.783
0.808
0.790
0.754
0.662
0.463
difficulty
metric
t=0
t=0.2
t=0.4
t=0.6
t=0.8
t=1.0
easy
skyline
0.487
0.487
0.487
0.488
0.481
0.471
Z
0.941
0.948
0.945
0.935
0.910
0.869
medium
skyline
0.151
0.134
0.116
0.108
0.102
0.093
Z
0.787
0.824
0.825
0.810
0.774
0.720
hard
skyline
0.007
0.005
0.006
0.006
0.005
0.003
Z
0.752
0.756
0.749
0.726
0.684
0.612
Part 2 — public-test potential
Pick the acceptance threshold k (max public-test failures allowed); every plot and table in this part updates to that k. k=0 is the strict “pass all public tests” potential.
skyline against temperature
skyline = (correct)/(accepted at threshold k), pooled. At k=0 this is the strict public-test skyline — much higher than the no-error skyline, and it tends to rise with temperature.
skyline by model and temperature
model
t=0
t=0.2
t=0.4
t=0.6
t=0.8
t=1.0
Llama-3 8B
0.706
0.694
0.676
0.694
0.686
0.723
Llama-3 8B Instruct
0.684
0.662
0.670
0.675
0.685
0.697
Llama-3.1 8B
0.670
0.662
0.673
0.686
0.707
0.734
Llama-3.1 8B Instruct
0.716
0.729
0.742
0.758
0.770
0.789
Llama-3.2 1B
0.579
0.590
0.563
0.501
0.481
0.412
Llama-3.2 1B Instruct
0.746
0.668
0.669
–
0.682
0.692
Qwen2.5-Coder 3B
0.766
0.772
0.777
0.789
0.793
0.825
Qwen2.5-Coder 3B Instruct
0.783
0.775
0.784
0.788
0.795
0.805
Qwen2.5-Coder 7B
0.783
0.799
0.808
0.816
0.825
0.844
Qwen2.5-Coder 7B Instruct
0.814
0.805
0.805
0.808
0.809
0.815
DeepSeek-Coder 6.7B
0.758
0.735
0.734
0.732
0.743
0.754
DeepSeek-Coder 6.7B Instruct
0.812
0.779
0.775
0.777
0.784
0.796
model
t=0
t=0.2
t=0.4
t=0.6
t=0.8
t=1.0
Llama-3 8B
0.342
0.334
0.308
0.289
0.249
0.195
Llama-3 8B Instruct
0.385
0.380
0.381
0.380
0.377
0.373
Llama-3.1 8B
0.326
0.313
0.306
0.290
0.258
0.207
Llama-3.1 8B Instruct
0.426
0.438
0.442
0.437
0.429
0.416
Llama-3.2 1B
0.120
0.097
0.083
0.056
0.041
0.019
Llama-3.2 1B Instruct
0.251
0.255
0.248
–
0.204
0.159
Qwen2.5-Coder 3B
0.497
0.489
0.488
0.487
0.463
0.425
Qwen2.5-Coder 3B Instruct
0.496
0.508
0.509
0.503
0.499
0.494
Qwen2.5-Coder 7B
0.551
0.554
0.553
0.541
0.519
0.479
Qwen2.5-Coder 7B Instruct
0.589
0.586
0.584
0.582
0.580
0.579
DeepSeek-Coder 6.7B
0.417
0.408
0.405
0.398
0.386
0.355
DeepSeek-Coder 6.7B Instruct
0.480
0.471
0.463
0.463
0.461
0.460
model
t=0
t=0.2
t=0.4
t=0.6
t=0.8
t=1.0
Llama-3 8B
0.152
0.146
0.131
0.113
0.083
0.045
Llama-3 8B Instruct
0.215
0.213
0.208
0.200
0.189
0.174
Llama-3.1 8B
0.143
0.140
0.132
0.115
0.087
0.050
Llama-3.1 8B Instruct
0.245
0.245
0.243
0.233
0.216
0.192
Llama-3.2 1B
0.029
0.023
0.018
0.012
0.008
0.003
Llama-3.2 1B Instruct
0.092
0.094
0.088
–
0.058
0.035
Qwen2.5-Coder 3B
0.291
0.291
0.285
0.270
0.231
0.162
Qwen2.5-Coder 3B Instruct
0.299
0.309
0.308
0.298
0.287
0.266
Qwen2.5-Coder 7B
0.351
0.353
0.348
0.327
0.287
0.210
Qwen2.5-Coder 7B Instruct
0.415
0.411
0.407
0.403
0.396
0.387
DeepSeek-Coder 6.7B
0.216
0.217
0.211
0.199
0.175
0.133
DeepSeek-Coder 6.7B Instruct
0.268
0.261
0.253
0.249
0.241
0.228
model
t=0
t=0.2
t=0.4
t=0.6
t=0.8
t=1.0
Llama-3 8B
0.110
0.105
0.093
0.077
0.053
0.025
Llama-3 8B Instruct
0.167
0.164
0.158
0.149
0.137
0.121
Llama-3.1 8B
0.104
0.102
0.094
0.079
0.056
0.028
Llama-3.1 8B Instruct
0.194
0.191
0.189
0.179
0.161
0.134
Llama-3.2 1B
0.016
0.013
0.010
0.006
0.004
0.001
Llama-3.2 1B Instruct
0.063
0.064
0.058
–
0.035
0.019
Qwen2.5-Coder 3B
0.224
0.225
0.219
0.203
0.165
0.102
Qwen2.5-Coder 3B Instruct
0.242
0.246
0.244
0.234
0.221
0.196
Qwen2.5-Coder 7B
0.283
0.285
0.278
0.256
0.213
0.135
Qwen2.5-Coder 7B Instruct
0.353
0.351
0.345
0.340
0.330
0.315
DeepSeek-Coder 6.7B
0.165
0.167
0.160
0.148
0.126
0.085
DeepSeek-Coder 6.7B Instruct
0.203
0.201
0.195
0.191
0.181
0.166
model
t=0
t=0.2
t=0.4
t=0.6
t=0.8
t=1.0
Llama-3 8B
0.108
0.103
0.091
0.076
0.051
0.024
Llama-3 8B Instruct
0.164
0.161
0.155
0.147
0.134
0.118
Llama-3.1 8B
0.103
0.101
0.093
0.077
0.055
0.027
Llama-3.1 8B Instruct
0.191
0.189
0.186
0.176
0.158
0.131
Llama-3.2 1B
0.015
0.012
0.010
0.006
0.004
0.001
Llama-3.2 1B Instruct
0.062
0.063
0.057
–
0.034
0.018
Qwen2.5-Coder 3B
0.221
0.222
0.215
0.199
0.161
0.098
Qwen2.5-Coder 3B Instruct
0.239
0.242
0.241
0.230
0.217
0.192
Qwen2.5-Coder 7B
0.280
0.281
0.274
0.252
0.209
0.131
Qwen2.5-Coder 7B Instruct
0.349
0.347
0.342
0.336
0.326
0.311
DeepSeek-Coder 6.7B
0.163
0.165
0.158
0.146
0.123
0.083
DeepSeek-Coder 6.7B Instruct
0.201
0.199
0.192
0.188
0.178
0.163
model
t=0
t=0.2
t=0.4
t=0.6
t=0.8
t=1.0
Llama-3 8B
0.108
0.103
0.091
0.075
0.051
0.024
Llama-3 8B Instruct
0.164
0.161
0.155
0.146
0.134
0.118
Llama-3.1 8B
0.102
0.100
0.093
0.077
0.054
0.027
Llama-3.1 8B Instruct
0.191
0.189
0.186
0.176
0.158
0.130
Llama-3.2 1B
0.015
0.012
0.010
0.006
0.004
0.001
Llama-3.2 1B Instruct
0.062
0.062
0.057
–
0.034
0.018
Qwen2.5-Coder 3B
0.221
0.221
0.215
0.199
0.161
0.098
Qwen2.5-Coder 3B Instruct
0.239
0.242
0.240
0.230
0.217
0.191
Qwen2.5-Coder 7B
0.279
0.281
0.274
0.251
0.208
0.131
Qwen2.5-Coder 7B Instruct
0.349
0.347
0.341
0.336
0.326
0.310
DeepSeek-Coder 6.7B
0.163
0.164
0.158
0.146
0.123
0.083
DeepSeek-Coder 6.7B Instruct
0.201
0.198
0.192
0.188
0.178
0.163
model
t=0
t=0.2
t=0.4
t=0.6
t=0.8
t=1.0
Llama-3 8B
0.108
0.103
0.091
0.075
0.051
0.024
Llama-3 8B Instruct
0.164
0.160
0.155
0.146
0.134
0.118
Llama-3.1 8B
0.102
0.100
0.093
0.077
0.054
0.027
Llama-3.1 8B Instruct
0.191
0.189
0.186
0.176
0.158
0.130
Llama-3.2 1B
0.015
0.012
0.010
0.006
0.004
0.001
Llama-3.2 1B Instruct
0.062
0.062
0.057
–
0.034
0.018
Qwen2.5-Coder 3B
0.220
0.221
0.215
0.199
0.161
0.098
Qwen2.5-Coder 3B Instruct
0.238
0.242
0.240
0.230
0.217
0.191
Qwen2.5-Coder 7B
0.279
0.281
0.274
0.251
0.208
0.131
Qwen2.5-Coder 7B Instruct
0.349
0.347
0.341
0.336
0.326
0.310
DeepSeek-Coder 6.7B
0.163
0.164
0.158
0.146
0.123
0.083
DeepSeek-Coder 6.7B Instruct
0.201
0.198
0.192
0.188
0.178
0.163
Z, the acceptance rate (normalizing constant)
Z = pooled proportion of rollouts accepted at threshold k (k=0: pass all public tests). Rises toward 1 as k grows.
by difficulty, choose a model
Public-test skyline and Z split by difficulty for a single model, at the selected k. Use both selectors; the plots and table update together.
skyline & Z by difficulty and temperature (selected model & k):