The question every hard set poses
You are four reps into a set of eight and you can feel where this is going. You could stop at eight. You could push to ten and find the rep that does not go up. How close to failure should you train, if what you want is muscle growth rather than a story about the set?
This is not a question about toughness. It is a question about price. Every rep you take past comfortable adds fatigue that has to be paid for by the sets after it, and by the session after that. If the last two reps buy you nothing, they are not free — they are expensive.
A systematic review with meta-analysis in Sports Medicine pooled the training studies that have actually tested this. Its answer is more specific than either side of the usual argument: the last rep to true failure appears to add nothing, and stopping well short of it costs you plainly.
Table of Contents
- What proximity to failure means and what got compared
- How close to failure should you train for muscle growth
- Grinding to failure did not win
- Stopping too early is the real mistake
- How close to failure should you train if you already lift
- When going to actual failure is worth it
- Where this sits in the wider evidence
- The limits of this evidence
- The practical answer
- Common Questions About Training to Failure
- How Pick It Up Helps
- Further Reading & References
What proximity to failure means and what got compared
Proximity to failure is how many reps you had left when you racked the bar. Most lifters know it as RIR, or reps in reserve: stop with two good reps still available and you trained at 2 RIR.
The thing being counted down to needs a definition, and this is where the research gets slippery. Momentary muscular failure is the strict version: the rep where you cannot complete the lift through a full range of motion without your form breaking down. It is the only definition that does not depend on someone’s judgement, because the set ends whether you like it or not.
Refalo and colleagues pooled 15 training studies that measured muscle growth, and their central methodological move was to stop treating all of those studies as if they asked the same question. They sorted them into three groups:
- Studies comparing training to momentary muscular failure against stopping short of it.
- Studies comparing training to some looser definition of failure — a prescribed rep count, a self-judged stopping point — against stopping short.
- Studies comparing different velocity-loss thresholds, where the set ends once your bar speed has dropped by a set percentage from your fastest rep.
That last one is worth understanding, because it turns out to carry the most useful information here. Velocity loss is a machine-measured stand-in for proximity to failure. Your reps slow down as you fatigue, so a set that ends at 20% velocity loss ends when your reps have slowed by about a fifth, and a set that ends at 40% goes considerably deeper.
The studies ran 6 to 14 weeks, training two to three times a week. Sizes ranged from 10 participants to 89.
How close to failure should you train for muscle growth
Here is the paper’s own summary of everything it found, arranged by how close to failure each condition trained.

The conceptual relationship the authors propose between proximity to failure and muscle growth. Each dot is how much muscle one training condition added, not a head-to-head comparison. From Figure 5 of Refalo et al. (2023), used under CC BY 4.0.
Those numbers on the vertical axis are effect sizes, which is how researchers express a change so that studies measuring different muscles with different machines can be pooled. They do not convert to inches on your arm. The shape is the point.
And the shape has two features. It climbs steeply on the left, then it stops climbing. The condition that stopped earliest sits at 0.20. The four conditions at the deep end — moderate velocity loss, high velocity loss, set failure and momentary failure — all land between 0.39 and 0.46, roughly double the first, and within that band going deeper stops buying anything. The final step to true failure, the dot on the far right at 0.41, is lower than the two conditions short of it.
One caution the authors are explicit about, and it matters for how much weight this curve can carry: the dots are not a direct comparison. Each one is how much muscle a given condition added, drawn from different studies with different participants. The curve is the authors’ proposal about the underlying relationship, not a measured contest. The direct comparisons are below, and they agree with it.
Grinding to failure did not win
The cleanest comparison in the paper is the one that isolates true failure: studies where one group trained to momentary muscular failure and another stopped short, with everything else matched.
Across those studies there was no detectable difference in muscle growth. The estimate favoured failure very slightly, and the uncertainty range around it runs from meaningfully worse to meaningfully better — which is the statistical way of saying the data cannot tell the two apart.
The velocity-loss studies land in the same place from a different direction. Ending sets at a greater than 25% speed drop produced no more muscle than ending them at 20 to 25%, and the estimate there was close enough to zero that it would be hard to build a case on. This is the comparison that matters most to anyone already lifting, because all six of those studies used trained participants.
There is one result in the paper that points the other way, and it deserves naming rather than burying. When all the studies were pooled together regardless of how they defined failure, training to failure came out ahead by a trivial margin that just cleared statistical significance. The authors then tested how much that result depended on a statistical assumption they had to make, and found it did not survive: nudge the assumption slightly and the finding disappears. They say plainly that it should be interpreted with caution. Treat it as noise rather than as a finding, which is what they do.
Two variables that might have explained the results away were tested and did not. Whether studies equated the total weight lifted between groups made no difference to the answer, so growth was not simply tracking who moved more tonnage. Load did not formally change the answer either, though the gap favouring failure was nearly twice as large in the studies using light loads (0.28) as in those using heavy ones (0.15). That difference was not statistically reliable, and the authors treat it as weak support for a specific idea rather than a finding: that stopping short matters more when the weight is light. It comes back below.
Stopping too early is the real mistake
The interesting half of this paper is not the failure result. It is the left edge of that curve.
The conditions that ended sets at less than a 20% velocity drop — stopping while the reps still looked fast — produced about half the muscle growth of the conditions that trained close to failure. That was the weakest result in the whole dataset, and it is the only estimate on the chart whose uncertainty range includes no growth at all.
So this is not a paper arguing that intensity of effort is overrated. It is a paper arguing that effort has a ceiling and a floor, and that most of the argument in gyms is about a region above the ceiling where nothing changes. You have to get close. You do not have to arrive.
The mechanism the authors propose for why the curve turns down is fatigue accounting. As a set approaches failure, your muscle fibres are recruited harder and exposed to more tension, which is the growth signal. But the fatigue from reaching failure carries into your next set and your next session, cutting the reps you can complete later. Past some point the extra tension in this set is paid for out of the tension in the following ones, and the ledger stops improving.
How close to failure should you train if you already lift
Here is the part that changes your week. Start with what this paper cannot give you, which is the number. It does not crown an optimum, and the authors say so directly: the proximity to failure that would maximise growth is still unclear, because the groups told to stop short never had their actual stopping point verified. In one included study, lifters squatting to a 40% velocity loss hit momentary failure about 56% of the time, so the “non-failure” condition was often failure.
What it does instead is bound the useful range at both ends, and that turns out to be the more practical result. There is a floor: stop while the reps are still moving fast and you give up about half the growth. There is a ceiling: the last rep to true failure adds nothing this literature can detect, while charging fatigue to every set behind it. Between those two lies a broad region where the data stops distinguishing anything.
Two to three reps in reserve is the sensible default, and it is worth being exact about the claim. This evidence does not show that 2 to 3 RIR beats 1 to 2 or 3 to 4, and anyone telling you it does is overselling it. What it shows is that the productive window is wide, that both edges of it are expensive, and that a target in the middle collects the growth at the lowest fatigue price. When evidence supports a range rather than a point, the middle of the range is the right place to sit.
That still leaves the question of which end of the set to worry about, and it is not the end most lifters worry about:
The error this data punishes is stopping too early, not going too deep. A set that ends while the bar is still moving fast is the one that cost people growth. Once the reps have visibly slowed, you are inside the productive window, and going deeper from there is optional rather than necessary.
Stop chasing the last rep on your heavy compounds. Whatever growth that rep buys is small enough that this literature struggles to detect it, and the fatigue it costs comes out of the sets behind it. A squat taken to the rep that fails is the most expensive set in your session and the one with the least evidence behind the cost.
Count your hard sets, not your failures. The authors’ own practical recommendation pairs a sufficient proximity to failure with adequate set volume, which they put at roughly 12 to 20 sets per muscle group per week. That figure is the more established of the two variables here — our review of how many sets per week for hypertrophy covers the dose-response evidence behind it. Getting the set count right is better understood than getting the last rep right.
Do not use soreness to judge whether a set was hard enough. Proximity to failure and next-day soreness are different things, and the evidence on what actually causes muscle soreness is that it tracks novelty more than it tracks effort.
When going to actual failure is worth it
The authors argue against treating this as a binary, and they list where they would bias a set toward genuine failure. Their reasoning is that failure costs fatigue, so spend it where the fatigue is cheap and the information is worth having:
- On simpler, lower-fatigue exercises. A single-joint movement over a multi-joint one, a machine over a free weight, anything that does not leave you breathing hard. A leg extension to failure costs you far less than a squat to failure.
- On the last set of an exercise or a muscle group. There is nothing left for the fatigue to damage.
- When that muscle group gets few sets — under five in a session, or trained only once or twice a week. If you are only giving it two sets, the case for making them count is stronger.
- When you are trained rather than new to lifting. Experienced lifters handle and recover from it better.
- When the load is light. People systematically underestimate how close to failure they are on lighter sets, so a light set nominally stopped short is often stopped very short.
That last point inverts on heavy work, which is the useful corollary. On a heavy set you can feel where failure is. On a set of twenty you cannot, so the stopping-short error is bigger and the argument for going to failure is stronger.
If you are compressing sessions and wondering where failure fits, our review of superset structures found that pairing two exercises for the same muscle cuts total volume, and this paper is why that trade is sometimes acceptable: it pushes you closer to failure on the second exercise.
Where this sits in the wider evidence
This study agrees with the consensus on its main question. Two earlier meta-analyses on failure versus non-failure training reached the same conclusion — most notably Grgić and colleagues in 2022, who concluded that training to failure does not appear to be required for gains in either strength or muscle size, and does not appear to harm them either. Refalo’s contribution is not overturning that; it is showing the result holds up when you stop mixing incompatible definitions of failure together, and adding the observation that the relationship bends rather than being flat.
One nuance in that earlier analysis cuts slightly against the tidy version, and it lands on exactly the group reading this. When Grgić’s team looked only at studies of people who already lifted, they did find a small but statistically significant advantage for training to failure for muscle growth. It was small, and it is a subgroup rather than a headline, but it is the one signal in this literature pointing at a real benefit for trained lifters specifically. Refalo’s analysis cannot check it: only two of its nine failure-versus-non-failure studies used trained lifters. Hold the “failure is not required” conclusion as well supported in general and slightly less settled for experienced lifters.
On the velocity-loss half, this paper does refine an earlier one. Hickmott and colleagues had reported that velocity losses above 25% beat those at or below 25% for muscle growth. Refalo’s team points out that this was driven mostly by comparisons against very low thresholds under 20%, and that once you separate 20 to 25% out as its own band, the advantage for going deeper disappears. Hickmott’s own paper reports the same split, so this is a difference in emphasis rather than in data. Both papers agree that stopping very early is worse for muscle growth; they read the region between 20% and 25% differently.
Hickmott’s analysis also found something this paper could not, because it measured strength as well as size: for gains in one-rep max, the lower velocity loss thresholds came out better, with the authors attributing that to accumulating less fatigue. If that holds, the last reps of a set are not just neutral for maximal strength but actively counterproductive, which points the same direction as the advice here for a different reason.
So the fair summary is that the failure-is-not-required finding is now reasonably well supported, and the shape of the curve below failure is a live question that this paper advances rather than settles.
The limits of this evidence
Most of the failure-versus-non-failure evidence is on untrained people. Seven of the nine studies in that comparison used participants who did not already lift. Only two used trained lifters. That is the weakest point in the paper for our purposes, and it is worth holding the failure result more loosely than the velocity-loss result, where all six studies used trained lifters.
Fifteen studies is a thin base. The authors say their subgroup analyses are probably underpowered and that they cannot rule out a different answer emerging from a larger literature.
Nobody knew how close the non-failure groups really got. This is the limitation the authors return to most often, and it caps what the whole literature can conclude. Until studies verify and report the proximity to failure actually reached, the precise shape of that curve is guesswork built on reasonable inference.
The methods reporting was uneven. No included study concealed which group participants were assigned to, none monitored what participants did outside the study, and only one clearly stated whether it analysed everyone it enrolled rather than only those who finished.
This is about muscle size, not strength. Every outcome pooled here is a measure of muscle growth. The paper does not tell you whether training to failure matters for a one-rep max, and there is reason to think strength and size behave differently — our plateau guide covers where they diverge.
Advanced set structures were excluded on purpose. Rest-pause and cluster sets were screened out, so none of this speaks to them.
Two of the five authors earn income as writers and practitioners in the fitness industry, which the paper discloses. No funding supported the review. Neither fact points in the direction of the result they found, which is a null one.
The practical answer
- Take most sets to two or three reps in reserve. Close enough that the reps are clearly slowing, short of the one that fails. That is the middle of the range this data supports and the cheapest place to sit in it.
- Stop grinding the last rep on heavy compound lifts. No detectable growth benefit, and a real fatigue cost to the sets behind it.
- Do not stop while the reps still feel fast. This is the error the data punishes: roughly half the growth of the conditions that trained close in.
- Spend your failure sets where they are cheap. Machines and single-joint work, and the last set of a muscle group rather than the first.
- Go to failure more readily on light, high-rep sets. You are worse at judging proximity to failure there, so nominally stopping short often means stopping far short.
- Get the set count right first. Roughly 12 to 20 sets per muscle per week is better-established than any RIR target, and it is the variable to fix before you optimise the last rep.
Common Questions About Training to Failure
Do you have to train to failure to build muscle? No. In the studies that compared training to true momentary failure against stopping short, there was no detectable difference in muscle growth. That result holds in this analysis and in earlier meta-analyses on the same question.
How many reps in reserve should I leave? Two to three on most sets. This paper cannot prove that figure is optimal, and it says as much — the growth-maximising proximity to failure is still unknown, because the groups told to stop short were never monitored. What it does establish is a floor and a ceiling: stopping well short costs real growth, and going all the way to failure adds nothing detectable. Two to three reps in reserve sits between those, which is the most this evidence honestly supports.
Is it bad to train to failure? Nothing here says failure damages your results. It says failure does not improve them while costing fatigue that your later sets need. On a low-fatigue exercise, or on your last set, that cost is small enough not to worry about.
Does stopping short mean I can do more sets? That is the trade this paper implies, and it is why the fatigue argument matters. The authors’ recommendation pairs a close-but-not-failure stopping point with 12 to 20 sets per muscle per week, and the volume evidence is the stronger half of that pairing.
Does this apply to strength training as well as muscle growth? Not from this paper — every outcome pooled here is a measure of muscle size. The earlier Grgić analysis did cover strength, and found no significant difference between training to failure and stopping short there either, so the two lines of evidence point the same way.
What about training light versus heavy? Whether studies used heavy or light loads did not change the failure answer overall. But the authors note that people misjudge proximity to failure more on lighter loads, so a light set is the one where you should be more willing to go all the way.
How Pick It Up Helps
The awkward thing about this research is that it hands you a target you cannot easily see. “Close to failure, not at it” is the right instruction and a hard one to follow, because judging how many reps you have left is a skill, and it degrades exactly when you need it — deep in a session, on a set that already hurts.
Pick It Up gives every set a target weight, a target rep count, and a target difficulty on a simplified RPE and RIR scale, so the intended proximity to failure is written down before you start rather than negotiated with yourself mid-set. Most sets are targeted at “medium”, which is our way of saying two to three reps in reserve — and this research is a good part of why.
That default is aimed deliberately between the two findings above. At two or three reps left the bar has slowed and the set is plainly hard, which puts it on the near-failure side of the floor where stopping early cost people half their growth. It also stops short of the failure end, where the extra rep bought nothing measurable and billed the rest of the session for it. And it matches what the authors themselves recommend: that most sets be taken to a close proximity-to-failure to limit accumulated fatigue, with genuine failure held back and decided on safety grounds. “Medium” is that recommendation turned into a number you can act on.
What we will not claim is that two to three is provably the best number. The evidence supports a window, not a point, and this is the middle of the window rather than a figure anyone has demonstrated to be optimal.
Then the app updates the next set’s targets from the weight, reps and difficulty you actually reported. A set that came in harder than planned changes what the following set asks for, which is the fatigue accounting this paper says matters, done for you.
Start a free trial and have all the science taken care of for you, so you can get back to building a better you.
Further Reading & References
- Influence of Resistance Training Proximity-to-Failure on Skeletal Muscle Hypertrophy: A Systematic Review with Meta-analysis — Refalo MC, Helms ER, Trexler ET, Hamilton DL, Fyfe JJ. Sports Medicine, 2023; 53(3): 649–665. PubMed, and the full text is free at PMC. The meta-analysis this article summarises, and the source of every number in it. Published open access under CC BY 4.0, which is also the licence on Figure 5 above.
- Effects of resistance training performed to repetition failure or non-failure on muscular strength and hypertrophy: A systematic review and meta-analysis — Grgić J, Schoenfeld BJ, Orazem J, Sabol F. Journal of Sport and Health Science, 2022; 11(2): 202–211. PubMed. Cited for where the consensus on failure training sits, for its strength finding, and for its trained-lifter subgroup result. Its abstract was read, not its full text.
- The Effect of Load and Volume Autoregulation on Muscular Strength and Hypertrophy: A Systematic Review and Meta-Analysis — Hickmott LM, Chilibeck PD, Shaw KA, Butcher SJ. Sports Medicine - Open, 2022; 8(1): 9. PubMed. Cited once, as the velocity-loss analysis this paper refines. Characterised from Refalo’s account of it and its own abstract, not from a full read.
- How Many Sets Per Week for Hypertrophy: Where Returns Diminish — our summary of the volume dose-response evidence, and the basis for the 12 to 20 sets per week figure the authors recommend alongside proximity to failure.
- Do Supersets Build Muscle, or Just Save Time? — why pairing two exercises for the same muscle pushes you closer to failure, and what it costs.
- What Causes Delayed Onset Muscle Soreness — why soreness is a poor read on whether a set was hard enough.
- How to Break Through a Plateau in Weightlifting and Strength Training — where strength and muscle size stop responding to the same things.
