Three hits prove nothing: how much testing a stat needs
✍️ Alice・愛麗沙
Your three screenshots prove nothing. The outermost layer of the damage formula is roundup(2 × rand(min damage×0.95, max damage×1.05) × …) — the weapon already has a minimum and a maximum, and the roll pushes each of them a further 5% outward. A 900/1000 weapon actually samples from 855 to 1050, so the top hit is 22.8% above the bottom one. To tell a 3% stat difference apart inside that noise you need roughly 60 hits per arm; for 1%, about 550. And if you have not separated crits from non-crits in your log, multiply the requirement by thirty-five. This is not an argument against testing. It is an argument for testing in a way that produces an answer.
Three layers of noise
From the outside in.
Layer one: the formula’s own roll. rand(min damage×0.95, max damage×1.05). Note the asymmetry — the two coefficients stretch the interval outward in both directions, 5% lower at the floor and 5% higher at the ceiling.
Layer two: the weapon’s own spread. The minimum and maximum printed on the weapon are already far apart before any roll is applied.
Layer three: crit. The worst of the three, and it gets its own arithmetic below.
Stack the first two. Take a 900/1000 weapon:
- lower sampling bound: 900 × 0.95 = 855
- upper sampling bound: 1000 × 1.05 = 1050
- highest ÷ lowest = 1.228
Same skill, same gear, same dummy, two hits, 22.8% apart. The stat you are trying to detect may be worth 2%.
How many hits is enough
Treat the outcome as uniform across 855 to 1050. The mean is 952.5, the standard deviation is 195 ÷ √12 = 56.3, and the coefficient of variation is 5.91%.
The sample size for comparing two means (95% confidence, 80% power) is:
hits per arm ≈ 2 × (1.96 + 0.84)² × (coefficient of variation ÷ difference you want to detect)²
Substituting:
- to detect 3%: 2 × 7.84 × (5.91 ÷ 3)² = about 61 hits (122 across both arms)
- to detect 1%: 2 × 7.84 × (5.91 ÷ 1)² = about 548 hits (roughly 1,100 in total)
- to detect 0.3% — the gap between two divisors, for instance — about 6,000 hits per arm
That last line is the honest bad news: some differences cannot be measured at home at all. Knowing that you cannot measure something is worth far more than measuring three hits and believing the result.
Crit multiplies the sample requirement by thirty-five
A crit is not “slightly more damage.” It splits your distribution into two separate lumps.
Take a 30% crit rate and a 2.0 crit multiplier. Mixed together, the mean multiplier is 1.3 and the standard deviation is √(0.3 × 0.7 × 1²) = 0.458, giving a coefficient of variation of 35.2%.
Sample size scales with the square of the coefficient of variation:
(35.2 ÷ 5.91)² ≈ 35.5
So the 61 hits from the previous section become 2,160.
Which means the first rule of stat testing is not “hit it more times.” It is pull the crits out and tabulate them separately. Split into two lumps, each lump’s coefficient of variation drops back to 5.91%, and the requirement drops straight back to the low sixties. It is the cheapest step available and the largest single win.
A historical error worth remembering
Early versions of the community formula wrote the outermost coefficient as 1.99. Later versions corrected it to 2, and the author explained why: 1.99 was an artefact of miscalculating the training dummy’s defence.
The lesson is not the 0.5% difference. It is this: when your measuring apparatus — the dummy — carries a variable you have not subtracted, you will measure a constant that looks entirely real, is wrong, and gets copied around for years.
Whether the dummy has defence, whether it has crit resistance, whether every dummy is identical — all of that has to be settled before the first hit, not after the last one.
What a test that counts actually looks like
- Same dummy, same skill, same standing position. Position changes how many hit instances land, and hit count changes everything downstream.
- At least 60 hits per arm. Fewer is defensible if you only want to know whether something does anything at all; if you are comparing two close options, plan for 500 or more.
- Log crits separately. This is the step that divides your sample requirement by thirty-five.
- Change exactly one variable. The cleanest experiment in the Korean record — swapping smash enhancement against multi-hit enhancement with the total held fixed at 5,112 — could conclude “identical damage to the last digit” only because nothing else moved.
- Check first whether the stat is worth testing at all. What 100 points buys is already tabulated:
| Stat | Divisor | Per +100 (from zero) | Notes |
|---|---|---|---|
| Multi-hit | 8500 | ≈ +1.18% | Shares one pool with its twin stat — only the sum matters |
| Smash | 8500 | ≈ +1.18% | Shares one pool with its twin stat — only the sum matters |
| AoE | 8500 | ≈ +1.18% | Only applies against 2+ targets |
| Skill power | 8500 | ≈ +1.18% | Multiplies the whole damage-increase bucket |
| Combo | 17500 | ≈ +0.57% | Also scaled by combo tier/4; full value only at 100+ combo |
| Crit | 10000 | ≈ +1.00% | Closed form: expected damage = 1 + crit/10000 |
| Extra hit | 13000 | ≈ +0.77% | Expected value of one extra 100% hit |
| Defence | 14828 | — | Mitigation = 1 ÷ (1 + DEF/14828), never reaches zero |
Divisors are KR community-model constants (identical across two independent implementations), not official; the per-+100 column is a marginal value from zero — returns diminish as a bucket fills.Stat efficiency
If the thing you want to test is worth 0.4% on paper, do not test it. The sample size you would need is an order of magnitude past your patience.
What to do with this
Do not argue with screenshots. Two screenshots can sit 22.8% apart on their own, and that gap has nothing to do with the stat under discussion.
Do not trust “it felt smoother after I swapped.” Against a 22.8% swing, that sentence carries no information.
Either sample enough, or say plainly that you did not. Every article here tries to mark which numbers are measured, which are derived, and which the original author explicitly told you to measure yourself — that labelling is the whole difference.
What could overturn this. The ±5% outer interval comes from the top-level expression of the Korean community model, a line marked high confidence — but no operator has ever published a combat formula, in any region. Every sample size above rests on the assumption that damage is roughly uniform inside the interval; if the real distribution is normal or triangular, the standard deviation is smaller than 195 ÷ √12 and the requirement falls with it (the direction of the conclusion holds, the numbers just get friendlier). If anyone accumulates a single-skill damage log large enough to draw the shape of that distribution, these figures get replaced with the measured ones.
