Independent work · github.com/AmroAbujabal/resnet-cifar-repro
Independent reproduction, not peer reviewed. Every number on this page is read from the results file produced by the linked code.
We reproduce the CIFAR-10 results of He et al. [1] for ResNet-20 and ResNet-56 from scratch, and obtain 8.39 ± 0.31% and 7.45 ± 0.69% test error over three seeds against the paper's 8.75% and 6.97%, both inside a pre-registered ±0.5% tolerance. We then run a controlled 2×2 comparing the original residual block against the full pre-activation block of He et al. [2] across CIFAR-10 and CIFAR-100, holding seeds, schedule, augmentation and parameter count fixed. Pre-activation produces no measurable change in test error at this depth: the differences are −0.23% and +0.22%, each well under the combined standard deviation, and they point in opposite directions. We report this as a null result and argue it is the expected outcome, since the advantage documented in [2] emerges at far greater depth.
Two claims are separable here, and only the first is a reproduction.
The first is that the CIFAR-10 error rates in Table 6 of [1] can be recovered from the paper's description alone. That is a reproduction claim, and it is tested against a tolerance fixed before any model was trained.
The second is an extension. Having a working baseline makes it cheap to ask whether the pre-activation block of [2] improves on the original at a depth where the original still trains well. That question has no paper number to match, so nothing in Section 4 is a reproduction claim. It is a controlled internal comparison.
Everything was written test-first against the paper text. The reference implementation of Idelbayev [3] was consulted for the Option A shortcut convention but not copied. The ImageNet-stem ResNet in torchvision is a different architecture and was deliberately not used.
The CIFAR architecture is the one specified in Section 4.2 of [1]. Inputs are 32×32 with per-pixel mean subtracted. A 3×3 convolution producing 16 channels is followed by 6n layers arranged as three stages of n residual blocks at 16, 32 and 64 channels, with stride 2 at each stage change, then global average pooling and a fully connected layer. Total depth is 6n+2. We use n=3 (ResNet-20) and n=9 (ResNet-56).
Shortcuts use Option A: when a stage changes dimension, the shortcut subsamples spatially with stride 2 and zero-pads the new channels, so the network stays entirely free of projection parameters. Parameter counts match Table 6 of [1] exactly, which is checked in the test suite rather than by inspection.
SGD with momentum 0.9, weight decay 1e-4 applied to all parameters, batch size 128, and 64,000 iterations on the full 50,000 image training split. The learning rate starts at 0.1 and is divided by 10 at iterations 32,000 and 48,000. Augmentation is 4-pixel padding, a random 32×32 crop and a random horizontal flip, with no colour augmentation. Evaluation is single-view on the 10,000 image test split.
The training split deserves to be stated rather than assumed, because holding out 5,000 images or not is worth roughly the size of the effects discussed here. Section 4.2 of [1] describes CIFAR-10 as “50k training images and 10k testing images” and says “we present experiments trained on the training set and evaluated on the test set”; the smaller split appears once, and only as the source of a hyperparameter — training terminates at 64k iterations, “which is determined on a 45k/5k train/val split”. So the paper's Table 6 models are trained on all 50,000 images, with 45k/5k used to fix the schedule length beforehand. This work follows the same two-stage arrangement: the 45k/5k split exists in the pipeline and is what the schedule was checked against, and every number reported here comes from a model trained on the full 50,000. Neither side trains on fewer images than the other, so the comparison in Section 3 is not carrying a data advantage in either direction.
Every configuration runs seeds 0, 1 and 2, and every reported figure is a mean with a sample standard deviation over those three runs. Each run appends one row to a single results file at completion, and every number on this page is read from that file. No result was transcribed by hand.
The definition of done for the reproduction was set in advance: the mean over at least three seeds must fall within 0.5 percentage points of the paper's error.
Table 1 CIFAR-10 test error against He et al. [1] Table 6. Three seeds per model, mean and sample standard deviation.
| Model | Params | Paper | This work | Δ | Per-seed error |
|---|---|---|---|---|---|
| ResNet-20 | 269,722 | 8.75% | 8.39 ± 0.31% | −0.36 | 8.30 / 8.13 / 8.73 |
| ResNet-56 | 853,018 | 6.97% | 7.45 ± 0.69% | +0.48 | 6.91 / 7.21 / 8.22 |
Both models meet the criterion. ResNet-20 is comfortably inside, landing slightly better than the paper. ResNet-56 passes by a margin of 0.02 percentage points, which is small enough that it deserves stating rather than rounding away.
The cause is a single run. Seeds 0 and 1 give 6.91% and 7.21%, straddling the paper's 6.97% closely, and on those alone the match would look tight. Seed 2 gives 8.22%, which is 1.31 points worse than seed 0 and well outside the pair. That one trajectory lifts the mean and doubles the standard deviation.
It was also the one run worth explaining and the one run with no record of it, so it was run again: same configuration, same seed label, under the deterministic setup and the per-interval logging of Section 2. It came back at 7.57%, which is 0.65 points away from the result it is nominally a rerun of, with a final train error of 0.03%.
The first thing that follows is uncomfortable. Under the original setup a seed label did not identify a run. Those numbers were produced with the cuDNN autotuner free to choose a different convolution algorithm per launch, so “seed 2” names a distribution and 8.22% is a draw from it that cannot be recovered on demand. The rerun is a fresh draw and not a replay, which is why it is reported beside the reproduction rather than inside it: Table 1 still stands on the original three runs, and a fourth measurement is not a fourth seed.
The second is that the excursion is unlikely to be an optimisation failure. The rerun fits its training set essentially completely — 0.03% train error against 7.57% test — which is the same fully fitted state as every cell in this work that measured train error at all. A network that reaches zero training error did not fail to optimise. This does not diagnose the original 8.22% run, whose train error was never measured and now never can be, but the reading that depth 56 is where optimisation starts to hurt finds no support here, and what is left is a generalisation gap that happened to be wide on one trajectory.
The honest summary is that the reproduction succeeds on the stated criterion, and that ResNet-56's interval is wide enough to overlap the paper's value rather than pin it.
He et al. [2] rearrange the residual block so that batch normalisation and ReLU precede each convolution, leaving the shortcut as a clean identity path. We compare that block against the original across two datasets, holding the seed set, schedule, augmentation and parameter count fixed.
The parameter budget is not approximately equal, it is identical. Moving the stem normalisation to a head normalisation on 64 channels costs 96 parameters, and the two dimension-changing blocks return exactly 96 by normalising input rather than output channels. Both variants have 853,018 parameters at n=9.
One detail decides whether this is the block from [2] at all. At a dimension-changing block the shared normalisation and activation are applied before the branch splits, so the Option A shortcut subsamples the activated signal rather than the raw input. Building it the other way trains without complaint and looks healthy in the loss curve while quietly being a different architecture, so the distinction is pinned by a test.
Table 2 The 2×2. Three seeds per cell, twelve runs total. CIFAR-100 has no baseline in [1], so those two rows are an internal comparison and not a reproduction claim. The original CIFAR-10 train error is marked rerun because none of that row’s three runs measured it; it comes from the single rerun described in Section 3 and does not enter that row’s test error.
| Block | Dataset | Test error | Per-seed error | Train error |
|---|---|---|---|---|
| Original | CIFAR-10 | 7.45 ± 0.69% | 6.91 / 7.21 / 8.22 | 0.03% (rerun) |
| Pre-activation | CIFAR-10 | 7.22 ± 0.29% | 7.08 / 7.03 / 7.55 | 0.02 to 0.03% |
| Original | CIFAR-100 | 29.75 ± 0.42% | 29.33 / 30.16 / 29.77 | 0.51 to 0.64% |
| Pre-activation | CIFAR-100 | 29.97 ± 0.37% | 29.96 / 30.35 / 29.61 | 0.48 to 0.49% |
Finding. At depth 56, the pre-activation block produces no measurable change in test error. On CIFAR-10 it is 0.23 points better, on CIFAR-100 it is 0.22 points worse. The combined standard deviations are 0.74 and 0.56, so each difference is 0.30 and 0.40 of the noise.
The sign reversal is the substantive part. A real effect of that magnitude would not flip direction between two datasets run under an identical protocol. Two differences that are small individually and contradictory jointly are better described as an absence of signal than as a weak signal.
This agrees with [2], which is the point of running it. The comparison in that paper at depth 110 on CIFAR-10 is close to even between the two blocks; the advantage that motivates pre-activation appears at 1001 layers, where the identity path's effect on gradient flow stops being marginal. Depth 56 is well below that regime, so a null is what the source paper predicts. The experiment confirms the boundary of the claim rather than the claim itself.
One asymmetry survives, and it is deliberately reported as an observation rather than a result. Pre-activation has the smaller standard deviation in both panels, 0.29 against 0.69 on CIFAR-10 and 0.37 against 0.42 on CIFAR-100, and the widest single excursion in the whole matrix is the original block's CIFAR-10 seed 2 — though the rerun of that seed in Section 3 fits its training set completely, so if the spread has a cause it is not a failure to optimise. Comparing dispersions from three samples per cell is weak evidence, and no further seeds were run, because adding seeds until a difference clears a threshold tests the experimenter rather than the network.
The CIFAR-100 cells fit their training data essentially completely. Train error settles between 0.48% and 0.64% across both blocks while test error stays near 30%, so what remains is generalisation gap and not underfitting. Both blocks reach that same fully fitted state, which is consistent with the null: they are not separated by optimisation capacity at this depth.
The shape is the familiar one from [1]: a long noisy plateau at the initial learning rate, a step change when it drops, then a short flat tail. It is worth noting that the second decay contributes almost nothing here, which suggests the schedule could be shortened without much cost, though that was not tested.
The rerun of Section 3 is the only run in this work with a train curve as well as a test one, and it separates the two effects the schedule produces. Both errors fall at the first decay, but not by the same kind of amount: test error moves from 12.90% to 7.80% and then flattens, while train error goes from 7.17% to 0.29% and keeps going to 0.03%. After the first decay the network is finished learning the training set and everything remaining is generalisation gap, which is the same state the CIFAR-100 cells reach and the reason the seed 2 excursion is not read here as a failure to optimise.
The repository holds the model, the training loop, the configuration for each of the five cells, the notebook that runs them, and the results file every number here is read from. The test suite covers the data pipeline, the metric, the model builder, an overfitting gate and the parameter counts, and it runs before any training does, so a typo fails in seconds instead of two hours into a GPU run.
Total cost was about 28 GPU-hours of training across 16 runs. Each configuration ran as one job, because a job killed at the session limit loses its output entirely.
Figures and every statistic on this page are written into it by scripts/build_site.py from results.csv and the per-run logs in the linked repository. Reproduction target and seed count were fixed before training began.