Abstract:
Title: Framing Resistance Training Program Design as an Optimum-Conditions Problem: A Formal Framework, Existence Theorem, and Evidence-Based Model for Endurance, Hypertrophy, Strength, and Power.
Background: Resistance training recommendations are frequently derived from expert opinion, mechanistic hypotheses, tradition, or partial readings of the research, and program critique is often binary ("it works" or "it does not") rather than relative. The article argues that "best" can be defined objectively as the highest expected value, the product of reliability (the frequency of positive outcomes) and effect size, and that program design differs structurally from intervention selection: every modifiable variable is set in every program, whether deliberately or by default, so the problem is one of specification rather than selection.
Objective: To propose a decision-theoretic framework that models resistance training program design as an optimum-conditions (specification) optimization problem over an outcome surface; to derive, as a theorem, that a single best model exists; and to present the Evidence-based Resistance Training Model (EBRTM) as the current estimate of the best region.
Eligibility criteria: This is a conceptual and mathematical framework paper rather than a systematic review. The proposed model assumes a bounded set of modifiable acute variables (load, repetition range, tempo, proximity to failure, rest, circuit training, sets per muscle group, frequency, set strategies, periodization, range of motion, exercise order, and exercise selection), a training goal chosen prior to optimization (endurance, hypertrophy, strength, or power), and a binding recovery constraint, with expected value estimates derived preferentially from peer-reviewed comparative research.
Information sources: The framework is grounded in expected value theory and the optimum-conditions problem formalized by response surface methodology (Box and Wilson, 1951), and integrates concepts from comparative effectiveness research, vote counting for evidence synthesis, and structured practitioner-level experimentation. The referenced model (EBRTM) was developed from systematic reviews of each modifiable acute variable, integrating findings from approximately 1,500 peer-reviewed studies with a focus on comparative research.
Risk of bias: No formal risk-of-bias tool was applied, as this is not a systematic review. The article instead argues for total-evidence inclusion over levels-of-evidence exclusion, sorting studies into comparable subgroups, extracting effect directions, and synthesizing trends by vote counting, with meta-analysis reserved for cases in which the direction of effect is unclear or a magnitude estimate could change a recommendation.
Results: The paper specifies twelve axioms, including probabilistic outcomes, expected value as the definition of "best," modifiable variables only, variable interactions, a recovery (rather than time) constraint, the optimum as a region represented by ranges, the objective as a surface rather than a sum, and goal-relative optimization. From these axioms, the paper derives an existence theorem: a bounded set of recoverable programs, each with a bounded expected outcome, must contain a maximum, and every program within the margin of equivalence of that maximum forms the best region, V* = { v : R(v) ≤ Rmax and E(v) ≥ E(v*) − e }. The solution is therefore a single model carrying recommended ranges, within which many programs tie for best and preference operates at no cost to outcomes. Comparison with the companion rehabilitation framework demonstrates that the two fields are different types of optimization problems (selection versus specification) sharing one method, and that each problem type nests as a subproblem within the other. The presented model (EBRTM) demonstrates that the four training goals share the majority of their programming, differing in a small set of goal-dependent variables (primary variable progressed, load emphasis, proximity to failure, tempo, exercise selection priorities, and advanced set strategies), and that comprehensive review revealed a variable omitted from traditional acute variable tables (proximity to failure).
Limitations: The paper acknowledges that comparative research can only evaluate the settings researchers have tested, so the estimated best region may be a local rather than global optimum; structured, one-variable-at-a-time experimentation with objective outcome measures is proposed to test settings beyond the current recommendations. A small number of recommendations rest on trends observed across reviews rather than dedicated systematic reviews and are labeled provisional estimates.
Conclusions: The article argues that a single best resistance training model exists, in the sense of a best region of an outcome surface derived from stated axioms, and that the model describing it can be estimated from comprehensive synthesis of comparative research and refined continuously by new evidence, objective outcome measurement, and deliberate experimentation. It proposes a testable framework in which critiques must identify a rejected axiom or misread evidence, and identifies the method (define best as expected value, identify the optimization problem type, state structural axioms, derive existence, initialize from comparative research) as generalizable to other outcome-comparable fields, including medical specialties.
Registration: Not registered (conceptual framework article).
Best Resistance Training Approach
A Formal Proof That a Best Resistance Training Model Exists, and the Evidence-Based Model That Estimates It
By Dr. Brent Brookbush, DPT, MS, BS, CPT, HMS, SPC, IMT
Introduction
A previous article defined the problem. In "Best Physical Therapy Approach (Physical Medicine): Evidence-Based Treatment Selection ," we asserted that choosing treatments is not a matter of opinion, convention, professional preference, or allegiance to a school of thought. "Best" is a measurable quantity. "Best" can be defined as "highest expected value," which is found by multiplying how often something works (reliability) by how much it works (effect size). Defining best as highest expected value allows us to compare techniques on the same scale, relative to one another. A question that sounds philosophical, "what is the best approach?" becomes a question that can be answered with an objective value.
The thought experiment that inspired that article. Imagine placing every physical rehabilitation technique ever conceived in a pile on a table: every modality, manual technique, and exercise from every profession. Which techniques do you pick? Remember, you can only pick as many techniques as will fit in a single session, so every technique you choose is a technique excluded. Most techniques that are currently used in practice help at least a little (survivorship bias). So, the question is not "what works?", but instead "what works best?" The pile is large, and the session is short. So, the best session is the combination of picks with the highest total expected value that can be completed within the session.
That thought experiment is an optimization problem. The term "optimization problem" is not a loose figure of speech. An optimization problem is any problem where the best possible result is needed given a set of rules and limits. Optimization problems are one of the oldest and most productive fields in mathematics. Mathematicians have been formally solving these problems since the 1690s, when some of the best minds in Europe competed to find the best shape of a curve to make a ball roll down to the bottom in the least amount of time. Euler and Lagrange soon turned puzzles like these into a general method for finding the best shape for anything smooth and continuous. A second branch of optimization problems was given serious study during World War II, when mathematicians were asked to find the best use of limited resources; for example, the best routes for convoys or the best way to load aircraft. Their work became the field now called operations research. The history of optimization problems matters for both this article and our previous article. When a problem can be recognized as an optimization problem, centuries of accumulated methods, proofs, and solved examples can be borrowed rather than reinvented. Recognizing that the problem is an optimization and what type of optimization problem it is, is most of the work, because the type of optimization problem determines which solutions apply. The "pile-on-the-table problem" from our previous article has the shape of a famous problem from operations research called the knapsack problem: many items, limited capacity, choose the combination with the greatest total value (1).
This article is a different optimization problem. Rehabilitation turned out to be a selection problem that was studied intently during WWII; that is, discrete choices under a hard constraint. Resistance training is also an optimization problem, but it is structurally different. It is not a selection problem. Which modifiable variables (a.k.a. acute variables) to use in a resistance training program is not an accurate description of the choice an exerciser must make. Instead, resistance training is a specification problem driven by a small set of coupled, modifiable variables. Every program must dictate load, repetitions, tempo, proximity to failure, rest intervals, sets, and training frequency. The central premise is this: every program sets every modifiable variable, regardless of whether the program explicitly made that choice. A lifter who has never heard of "repetition tempo" still lifts at a tempo. A program written without intentional periodization simply has a periodization setting of "constant". You cannot skip a variable; you can only set it deliberately or by accident. This is the same problem a statistician and a chemist (Box and Wilson) formalized in 1951, when they asked how the temperature, pressure, and timing of a chemical process should be set to maximize its yield (2). The type of optimization problem doesn't have a catchy one-word name the way "the knapsack problem" does. Box and Wilson's own paper called it "the experimental attainment of optimum conditions," so the problem is usually called the optimum conditions problem or, more generically, process optimization. The method they invented to solve it does have a famous name: response surface methodology (RSM), which is still the standard term used in statistics and industrial engineering today.
An analogy to better understand the resistance training optimization problem. The following analogy describes the problem. You are standing in front of a machine covered in dials, one dial per variable. A load dial. A reps dial. A tempo dial, a rest dial, a sets dial, a frequency dial, and so on down the panel. The dials make it possible for the machine to optimize production of four distinct products (only one at a time): endurance, hypertrophy, strength, and power. Pick the product you want most, then set the dials to produce the best possible product. Three things about this machine define the entire problem. First, no dial can be left unset. Every dial is set to some number the moment the machine is turned on, including all of the dials you did not touch; if you did not change a dial, it will simply continue to be set at that number. Second, the dials are not independent. Turning one dial may change the best position of another, the way turning up the volume on a stereo may require you to adjust the bass. Third, the machine's production is somewhat forgiving. The best position of most dials is a range, and small differences within that range barely move the needle. The rest of this article asserts that the best range for every dial, for every product, can be found. A model built by aggregating all of the comparative research for each modifiable variable is the most accurate model that can currently be developed to describe the optimal settings (the modifiable variables) for each product (each resistance training goal).
What is borrowed, and what is new. It is worth being precise about the innovations claimed in this article. The mathematics is borrowed: expected value has been used to define rational choices in decision theory and economics since the 1940s, and the optimization methods described above have even longer histories. The contribution of this article, the previous article, and the models and recommendations that have followed is the novel application of these methods. The first innovation was applying expected value to define "best" in fields where "best" was still being decided by convention, professional preference, and allegiance to schools of thought. The importance of this step cannot be overstated. It changed "best approach" from a subjective judgment to an objectively measurable value, and a measurable value made the problem addressable by the scientific method, reducing the reliance on expert belief. The second innovation was recognizing that models would be required. Physical rehabilitation intervention selection and resistance training program design are too complex to solve one recommendation at a time; a model is required to both determine the best set of values and make those values practically useful. The third innovation was recognizing that these models are optimization problems. This realization further supported the idea that a "best approach" could be developed, and it helped make the problem solvable, because identifying the type of optimization problem determines which solutions apply. The fourth innovation was recognizing what kind of research can give the models their values: research that prioritizes outcomes over mechanistic hypotheses, and specifically comparative research, because only comparative research can determine the relative efficacy of each option. The fifth innovation was recognizing that the likelihood of accuracy increases with the amount of research reviewed, and is highest with a complete and comprehensive review of the best available research. The applications of borrowed mathematics were the innovations that opened the door; however, the majority of the labor is the immense amount of research collected and synthesized to develop recommendations. To our knowledge, no one else has attempted this in a way that is comprehensive and complete. The result is models and recommendations built from many times the amount of relevant research used to develop previous models, with an unprecedented likelihood of accuracy and optimal efficacy.
Axioms of the Resistance Training Optimization Problem
Why start with axioms? An axiom is a foundational statement, a self-evident truth, or baseline principle accepted as true that does not need further proof to establish that it is true. It serves as a starting point for further reasoning, logic, or mathematical systems. The axioms in this article force every assumption into the open, where each assumption can be checked and challenged. Further, the axioms intend to define the problem completely, so the reader can confirm that the model solves the problem it claims to solve. A previous article addressed the "Best Physical Therapy Approach " and required its own set of twelve axioms. The resistance training problem is a different type of optimization problem, requiring a different set of axioms. The twelve axioms below are intended to be complete for this problem. Together, these axioms define the goal of the optimization, the variables that can be modified, the constraint that limits those variables, and the evidence used to determine the best settings. Note that the existence of a single best model is not asserted as an axiom. The existence of a single best model is derived from the axioms, as a theorem, at the end of the formal model section. Axioms shared with the previous article are labeled "shared with the previous article," and axioms unique to this problem are labeled "specific to resistance training"; the differences between the two sets of axioms are discussed in a later section.
Axiom 1: Outcomes are probabilistic. A program does not guarantee a result; results are probabilistic. Some people will improve more, some will improve less, a small number may not improve at all, and some may have incredible results that exceed expectations. Every recommendation in this model is the range most likely to get the best result for the greatest number of people (highest expected value - Axiom 2) (shared with the previous article).
Axiom 2: "Best" is highest expected value. Expected value is reliability (how often) multiplied by the magnitude of the effect (how much improvement). Defining best as highest expected value makes "best" an objectively measurable quantity rather than a subjective measure of preference. Modifiable variable range recommendations are intended to result in the highest expected value based on the relative effect on outcomes (shared with the previous article).
Axiom 3: Only modifiable variables belong in the model. The model only includes variables that can be directly modified within a program: load, repetitions, tempo, proximity to failure, rest, sets, frequency, exercise selection, exercise order, range of motion, set strategies, and periodization. Factors that cannot be directly modified, or cannot be modified enough to change outcomes, are excluded. For example, the acute spike in testosterone and growth hormone that follows a workout cannot be directly set by a practitioner, and research suggests it likely has little effect on long-term outcomes, so it is excluded. This axiom is different from the axiom used in the rehabilitation model. In physical rehabilitation, the choices were between all of the available techniques; in resistance training, the choices are set to a limited number of variables that can be directly modified within a program (specific to resistance training).
Axiom 4: The variables interact, and the interactions are part of the model. The variables are mostly independent, but not completely independent. The best setting of one variable can depend on the setting of another. For example, the best rest interval depends on the number of sets performed, the stability of an exercise may depend on the load, and the number of sets should be reduced when drop sets are introduced (Acute Variables: Set Strategies ). A complete model does not set each variable in isolation; a complete model states these dependencies. Identifying the interactions is what turns an impossibly large search (every combination of every setting) into a manageable search. Note that the interactions only become visible when all of the variables are reviewed, which is one reason a complete review is required (Axiom 11) (specific to resistance training).
Axiom 5: The binding constraint is recovery, or the limitations of the body's ability to adapt to a stimulus, not the amount of time in a session. In the rehabilitation model, the constraint was session length: only so many techniques fit in a session. In resistance training, the constraint is the body's capacity to adapt or recover. Training volume cannot be increased without limit, because past a certain point additional volume produces smaller improvements, and eventually worse improvements. The research on sets per muscle group demonstrates this directly: improvements increase up to approximately 3 sets per muscle group per session, increase less with the addition of a 4th and 5th set, improve little beyond that range, and may decline at 6 or more sets (specific to resistance training) (Acute Variables: Sets per Muscle Group ).
Axiom 6: The optimum is a region, and ranges represent it. For most variables, outcomes are similar across a span of settings. The best setting is therefore not a single number but a zone, and the model represents each zone as a range (for example, 3-8 RM, or 2-3 minutes of rest). This is a critical point: a recommendation stated as a range is a precise claim, not a vague claim. The claim is that outcomes are equivalent within the range and worse outside of the range. This representation of an optimum has precedent outside of exercise science: pharmaceutical manufacturing formally approves a "design space," a region of process settings within which any combination of settings has been demonstrated to produce an acceptable product, and movement within that region is not even classified as a change (15). This axiom is what allows the phrase "a single best model" to be accurate even though no single best number exists for every variable. The model is singular; the ranges are carried inside the model (specific to resistance training).
Axiom 7: The objective is a surface, not a sum. Every variable is set in every program, so the total expected outcome is a function of all of the settings together. Imagine a combination of modifiable variable ranges, a recommendation for all of the modifiable variables for a particular goal, as a point on a map. Now, consider the expected value of those recommendations to be the point's elevation. Every modifiable variable adjustment results in a new point, with a new elevation, and we could put very similar recommendations next to one another, so their elevations could be compared. If we mapped results this way, we could visually see which point (set of recommendations) resulted in the highest elevation (best outcome - highest expected value). This visualization is very similar to how this type of optimization problem is solved, and is referred to in this article as the "outcome surface." The optimization problem is finding the highest region on that surface, and is what was referred to earlier as the "response surface methodology (RSM)." Each modifiable variable, the dial on the machine from our analogy, sets the parameters for a point on our map. This objective is different from the objective in the rehabilitation model, which summed the expected values of the techniques selected. In rehabilitation, every technique chosen was a technique excluded; in resistance training, setting one variable does not "use up" resources needed by another; all variables must be set, either by choice or default (specific to resistance training).
Axiom 8: The goal is given, not chosen by the model. Optimization only has meaning relative to a goal. This model identifies the best program for a chosen goal (endurance, hypertrophy, strength, or power); it does not determine which goal an individual should pursue. Choosing the goal is the individual's job; recommending the best modifiable variable ranges is the model's job. Note that goals outside of resistance-training-specific adaptations (for example, weight loss, general fitness, or longevity) do not require separate models, because the research suggests that the choice of resistance training program has little effect on those outcomes, or that those long-term goals are best addressed by a specific resistance training adaptation (e.g., hypertrophy may be best for longevity, it likely does not matter for weight loss, and power is ideal for athletes) (specific to resistance training.)
Axiom 9: Values are initialized from comparative research. The reliability and effect size of each setting are estimated from peer-reviewed comparative research: studies that test two or more settings of a variable against each other using a reliable, objective outcome measure. Comparative research is required because it answers the question this model asks. A study without a comparison can demonstrate that a program works; it cannot demonstrate whether a better setting was available (shared with the previous article).
Axiom 10: Comparative findings are synthesized by vote counting, not pooled averaging. Studies differ in populations, doses, durations, and outcome measures. Rather than forcing these differences into a single averaged number, studies are first sorted into comparable subgroups, and then the direction of effect is counted within each subgroup: favors A, favors B, or no significant difference. Sorting first is what turns "mixed" research into a clear trend. Meta-analysis is not treated as a higher form of evidence; meta-analysis is a different tool, reserved for cases where the direction of effect is unclear or where a representative magnitude would improve a recommendation (shared with the previous article).
Axiom 11: Completeness raises the probability of finding the true optimum. A model built from more of the relevant comparative research has a higher probability of locating the true best region, because more evidence bears on every estimate, and because the interactions between variables (Axiom 4) only become visible when every variable has been reviewed. Note that accuracy is a probability, not a guarantee. A complete review does not guarantee the right answer for any single variable, and a smaller review could stumble onto the same answer. What a complete review guarantees is that no recommendation was wrong because of evidence that was never consulted, and that guarantee raises the expected accuracy of the whole model (shared with the previous article).
Axiom 12: Deliberate experimentation is required to confirm the optimum is global. Comparative research can only test the settings researchers chose to test. If a better setting exists outside every tested range, no review can find it. The current best region may therefore be a "local" optimum: the best of everything tried so far, rather than the best possible. Confirming that the optimum is global requires deliberately testing settings outside the current best region and updating the model based on reliable, objective outcomes. Within the model, autoregulated load adjustment is already a small version of this process on a single variable: performance in each session is used to nudge the load setting up or down. The full version, systematic experimentation aimed at the surface as a whole, is described in a later section. It should be noted that experimentation that overlaps with methods already represented in the research does not result in progress toward a better global maximum. This is another reason why completeness is important. Without completeness, ignorance could result in simply retesting tried methodologies (shared with the previous article).
Formal Optimization Model for Resistance Training
The previous article expressed the rehabilitation model in formal notation, and the same is done here for resistance training. Formal notation is not decoration. Writing the model as mathematics forces every term to be defined, exposes exactly what is being maximized and what is limiting the maximization, and allows the model to be checked, criticized, and improved. The notation below requires no advanced mathematics to read; every symbol is defined in plain language.
Let:
- v (version) = (v₁, v₂, ... vₙ): The program, written as a list of settings, one setting per modifiable variable (load, repetitions, tempo, proximity to failure, rest, sets, frequency, exercise selection, exercise order, range of motion, set strategies, and periodization).
- g (goal): The chosen training goal (endurance, hypertrophy, strength, or power). The goal is chosen before optimization begins (Axiom 8), so all terms below are defined relative to the chosen goal.
- E(v) (expected value of version): The expected outcome for the chosen goal when the program is set to v. This is the "outcome surface" described in Axiom 7: every possible program is a point, and E(v) is the elevation at that point.
- p and m (reliability and magnitude): For any comparison of settings, p is the reliability (how often a setting produces the better outcome) and m is the magnitude (how much improvement). Expected values (p × m) estimated from comparative research are the data used to construct E.
- R(v) (recovery cost of version): The recovery cost of the program; that is, the total training stress the program imposes, driven primarily by volume (sets, frequency, and proximity to failure).
- Rₘₐₓ (recovery max): Recoverable capacity; the maximum training stress from which the body can recover while continuing to adapt.
Objective Function: Maximize the Expected Outcome for the Chosen Goal
Maximize E(v)
- The optimization problem is to find the program v that produces the highest expected outcome for the chosen goal. The goal is chosen by the individual, not by the model (Axiom 8), and "highest expected outcome" is defined by expected value (Axioms 1 and 2).
Recovery Constraint
Subject to: R(v) ≤ Rₘₐₓ
- The program cannot impose more training stress than the body can recover from. This constraint replaces the session-time constraint of the rehabilitation model (Axiom 5). The sets-per-muscle-group research is the direct evidence of this constraint: expected outcomes stop improving, and may decline, when volume exceeds recoverable capacity.
The Objective is a Surface, Not a Sum
E(v) ≠ f₁(v₁) + f₂(v₂) + ... + fₙ(vₙ)
- The expected outcome of a program is not the sum of independent per-variable scores, because the variables interact (Axiom 4): the best rest interval depends on the number of sets, exercise stability interacts with load, and sets should be reduced when drop sets are introduced. This inequality is the formal difference between the two articles. The rehabilitation model maximized a sum of selected items under a time budget (a knapsack problem); this model maximizes a joint function of simultaneous settings under a recovery constraint (an optimum conditions problem).
The Solution is a Region, Not a Point
V* = every program v for which E(v) is statistically indistinguishable from the maximum
- Because outcomes are similar across a span of settings for most variables (Axiom 6), many programs tie for best, and the solution to the optimization is the set of all of them: the best region, written V*. The recommended range for any single variable is that variable's slice of the best region. This is why the model can be singular while its recommendations are ranges: V* is one region, every program inside V* is expected to produce equivalent outcomes, and every program outside V* is expected to produce worse outcomes (proved in the theorem at the end of this section).
Initialization and Refinement
- The surface E cannot be observed directly; it must be estimated. The estimates are initialized from peer-reviewed comparative research (Axiom 9), synthesized by vote counting within comparable subgroups (Axiom 10), and improved by completeness, because more comparative research raises the probability that the estimated best region contains the true best region (Axiom 11).
- The estimated best region is confirmed and expanded by deliberate experimentation: testing settings outside the current best region and updating the model from reliable, objective outcome measures (Axiom 12). Autoregulated load adjustment is this process operating continuously on a single variable.
The Complete Model
The pieces above can now be assembled into one complete statement. Two additional symbols are needed:
- v* (best version): Any single program that produces the maximum expected outcome for the chosen goal. Formally, v* is the program that satisfies R(v*) ≤ Rₘₐₓ and E(v*) ≥ E(v) for every other program v that also satisfies the recovery constraint.
- e (margin of equivalence): The largest difference in expected outcomes that research cannot distinguish from no difference. If two programs differ by less than e, the research cannot tell them apart, so the model treats them as equivalent.
The complete model, written as one formula:
- V* = { v : R(v) ≤ Rₘₐₓ and E(v) ≥ E(v*) − e }
The formula is read left to right as follows:
- V* = ...The best region (the solution to the optimization problem) is...
- { v : ...the set of every program v such that...
- R(v) ≤ Rₘₐₓ ...the program's recovery cost does not exceed recoverable capacity (Axiom 5)...
- and E(v) ≥ E(v*) − e } ...and the program's expected outcome is within the margin of equivalence of the best possible program (Axioms 1, 2, and 6).
Read in plain language: The solution to the resistance training optimization problem is the set of every recoverable program whose expected outcome is statistically indistinguishable from the best possible outcome. That set is the best region, V*. The recommended range for each modifiable variable is that variable's slice of V*, which is why the model's recommendations are ranges (Axiom 6), why many programs can tie for best while remaining inside one model (see the theorem below), and why every program outside V* is expected to produce worse outcomes. The remainder of this article presents the evidence used to estimate V* for each training goal, and the method for confirming that the estimated region is the true one (Axiom 12).
Theorem: A Single Best Model Exists
The sections above asserted axioms; this section derives the central claim of this article from those axioms. A theorem is a statement that must be true if the axioms it is derived from are true. The thesis of this series, "a single best model exists," is deliberately not one of this article's axioms. The thesis is a theorem, because it follows necessarily from the twelve axioms already stated. The derivation can be followed step by step.
- First, the problem is bounded. The model contains a finite list of modifiable variables, and each variable has a limited span of possible settings (Axiom 3). Every possible program is a combination of those settings, so the set of possible programs is enormous, but it is not unlimited.
- Second, every program has a value. The goal is chosen before optimization begins (Axiom 8), and every program has an expected outcome for that goal, E(v), determined by reliability and magnitude (Axioms 1, 2, and 7). Expected outcomes are also bounded; no program produces unlimited improvement.
- Third, the "ability to recover" constraint narrows the set without emptying it. Only programs the body can recover from are candidates for the best possible program (Axiom 5). Recoverable programs obviously exist, so the set of programs a body can recover from is smaller, but it is not empty.
- Fourth, a bounded set of recoverable programs, each with a bounded value, must contain a maximum. This is a standard result in mathematics: if the options are limited and every option has a value, at least one option has the highest value. Therefore, at least one best program (v*) exists.
- Fifth, the maximum is a region, not a point. Research cannot distinguish differences in outcomes smaller than the margin of equivalence, e (Axiom 6). Every recoverable program within e of the best program is therefore tied with it, and the solution is the set of all tied programs: the best region, V*, exactly as written in the complete model above. The region contains at least one program (v*), so the best region exists.
- Last, refinement converges toward the true region. The location of the best region is estimated from comparative research (Axiom 9), synthesized objectively (Axiom 10), improved by completeness (Axiom 11), and tested by deliberate experimentation (Axiom 12). Each addition of evidence raises the probability that the estimated region contains the true region, so the model improves rather than becoming outdated.
It is worth being precise about what this theorem does and does not establish. The theorem establishes that, if the twelve axioms are true, a best region of the outcome surface must exist, and the model that describes it is singular. A critic of this article's thesis must therefore identify which axiom they reject; rejecting the conclusion alone is not available, because the conclusion follows from the axioms. The theorem does not establish two other claims, and both are addressed in later sections. The theorem does not establish that the best region is one connected region rather than two distant regions of equal height; that claim is supported by probability, not proof, in the section "Chance of Multiple Best Solutions." The theorem also does not establish that the current model has located the true best region; that claim is empirical, and the evidence for it is the comparative research base and the refinement process (Axioms 9-12). This is a defining thesis of this article series; there may not be a single best number for every variable, but there is a single best model.

Section 1: There is A Best Approach
Outcomes are probabilistic, not deterministic.
A resistance training program does not determine a result; it changes the probability of a result (Axiom 1). Programs are often sold and taught deterministically, as if a given program will reliably produce a given outcome ("this program will add 10 pounds of muscle," "8-12 reps builds muscle, 1-5 reps builds strength"). In reality, the same program produces a spread of outcomes for the people who perform it. Research on individual responses demonstrates this directly: when large groups perform an identical program for the same number of weeks, changes in muscle size and strength range from no measurable improvement to improvements several times larger than the group average (3). Some of this variation reflects factors the model excludes because they cannot be modified (for example, genetics), some reflects factors outside the program (for example, sleep, stress, and nutrition), and some remains unexplained.
Probabilistic outcomes do not weaken the claim that a best program exists; probabilistic outcomes define what "best" means. If outcomes were deterministic, the best program would be the one program that works, and any variation between individuals would be proof of error. Because outcomes are probabilistic, the best program is the program with the highest expected value: the settings most likely to get the best result for the greatest number of people (Axioms 1 and 2). Every recommendation in this model is a statement of that form. A recommendation of 3-8 RM for strength does not promise any individual their best possible result; it asserts that, across a population, no other load range is expected to produce better strength outcomes.
Probabilistic outcomes also explain why an individual result cannot validate or refute a program. One lifter's exceptional progress on an unusual program is not evidence that the program is best, any more than one smoker living to 95 is evidence that smoking is safe. The lifter may be a high responder (an individual who improves more than average on almost any program), and a high responder may have progressed even more on a better program. The reverse is also true: one lifter's poor result on a well-supported program does not refute the recommendation; it most likely identifies a below-average responder (an individual who improves less than average on almost any program), or an unmeasured factor outside the program. Both errors are versions of the availability heuristic (the tendency to weight memorable individual results more heavily than measured frequencies). Determining what works best requires comparing outcomes across many people, which is exactly what comparative research does (Axiom 9), and adjusting for the individual requires objective measurement over time, not memorable anecdotes (Axiom 12).
Critique is not binary.
The efficacy of a modifiable variable recommendation does not compete with doing nothing. The efficacy of a recommendation competes with every other possible recommendation for that same variable (Axiom 2). When investigating the best repetition range, set strategy, or training method, the relevant questions are not binary (can be answered with a "yes" or "no"). Nearly every recommendation will contribute to some strength, hypertrophy, endurance, or power. The question that should be asked is relative; does this recommendation result in better outcomes than other recommendations? The intent of the model is not to determine "what works," but instead to determine "what works best."
Survivorship bias (the tendency for only successful examples to remain visible, because failures disappear from view) explains why most methods in current use "work." Methods that consistently produce no result fade out of popularity. Methods that consistently produce much worse results do not spread as readily. The programs, repetition ranges, and technique recommendations that remain in gyms and certifications, and that are promoted on social media, are the survivors, and nearly all of the survivors produce some improvements for some populations. This is why testimonial evidence is abundant for every method simultaneously, including methods that directly contradict one another. Survivorship almost guarantees that every surviving method will have testimonials to support it; however, survivorship bias does not guarantee promotion of "what works best."
Resistance training includes a factor that strengthens survivorship bias and makes "what works best" harder to determine without a scientific method; rehabilitation does not share this factor. For many outcome measures, novice lifters improve rapidly during the first weeks of resistance training, regardless of how the modifiable variables are set. Untrained individuals improve on nearly any program. A program does not need to be well designed to produce a great testimonial; a program only needs to be performed by a beginner. This makes "it worked for me" even weaker evidence in resistance training than in rehabilitation. The improvement is real, but the improvement is evidence that resistance training works; not evidence that any single program worked best. This rapid initial improvement on nearly any program is sometimes referred to as the "novice effect." Comparative research becomes essential for determining the best possible, modifiable variable recommendations, because anecdotal evidence, which is already a weak form of evidence, may be completely confounded by the novice effect. Further, research investigating more experienced exercisers is needed because they are likely to respond very differently.
Last, the most common debates in the industry may be resolved with the concept of "relative effectiveness." Debates are routinely framed as binary questions: machines versus free weights, full range of motion versus partial range of motion, sets to failure versus reps in reserve, high frequency versus low frequency. When these debates are framed as binary questions, both sides can claim that their option "works," and the debate cannot be resolved. When these debates are framed as relative questions, each debate becomes a comparison of recommendations for a single variable, and comparative research can resolve it: which recommendation produces the best outcome for a chosen goal (Axioms 2 and 9)? Every modifiable variable recommendation is treated this way in this article. The answer is never "yes" or "no"; the answer is a range, for a goal, supported by the direction of effect across the comparative research (Axioms 6 and 10). The debates ask "what works?"; the model asks "what works best?"
Unsupported Default Position Fallacy
Another common reasoning error must be addressed: the unsupported default position fallacy. This fallacy is the belief that identifying a flaw in an opposing argument automatically strengthens one's position. In short: "Proving you wrong makes me right." This is fallacious unless there are only two mutually exclusive options, and one of them must be correct. That is rarely, if ever, the case in resistance training, where there are many modifiable variables, each with a range of recommendations. When more than two recommendations are possible, or when multiple recommendations may be simultaneously flawed, each recommendation must be evaluated on its own merits. Demonstrating flaws in one recommendation does not excuse or validate the alternative; it still must be shown that the other recommendation is less flawed or has greater expected value (Axioms 2 and 9).
This fallacy is particularly rampant on social media, where "debunking" and "response" content are very popular, but often these posts are rewarded for the takedown and omit the comparison. A good example is seen in debates over the optimal rep range for hypertrophy. Research demonstrates that hypertrophy occurs following a wide range of loads and rep ranges, so it is often stated that "rep ranges do not matter." However, demonstrating that 8-12 repetitions per set is not uniquely effective does not demonstrate that all repetition ranges are equally effective. A reverse argument is equally fallacious: demonstrating that a rep range is effective does not prove that it is optimally effective. The best argument should be the argument best supported by the best evidence (research). In this case, the research suggests that a wide range of loads and reps are effective for hypertrophy; moderate loads and reps (8-12) are likely optimally effective, and other variables like "sets to failure" may be more influential (Axioms 2, 6, 9, and 10) (Hypertrophy Training: Evidence-based Model ).
A related fallacious argument is dismissing a recommendation because the explanation is flawed, despite an outcome occurring. Refuting a mechanistic hypothesis of effect may be reasonable, but an outcome is not dependent on a mechanistic hypothesis. If an outcome occurs, but the hypothesized mechanism of effect is proved wrong, what is needed is a new hypothesis for why an outcome occurred. The outcome itself cannot be denied. The reverse would also be true. If no outcome occurred, no matter how convincing the mechanistic hypothesis of effect, the mechanistic hypothesis cannot be used to imply an outcome that did not happen. What determines the efficacy of a recommendation is the outcomes, not the mechanistic hypothesis of effect. For example, it was believed that the stress of lengthened partials would result in more hypertrophy because the lengthened position resulted in more stress, tension, and stimulus to muscle fibers; however, the research did not demonstrate that lengthened partials were more effective than full range of motion (ROM) repetitions. However, adding lengthened partials to the end of a set to perform a few more repetitions, after full ROM repetitions to failure, may result in more hypertrophy, as this is closer to a "drop set," which is an effective strategy for increasing hypertrophy (See the articles: Range of Motion (ROM) and Hypertrophy: Systematic Review and Drop Sets: Comprehensive Systematic Review and Training Recommendations ).
It is likely that every recommendation and mechanistic hypothesis of effect in resistance training will require refinement, just as every scientific domain does. Exercise science is an evolving field. Identifying flaws is essential to progress, but doing so does not imply that an opposing position is automatically better. Ironically, many of the alternative hypotheses asserted on social media to oppose an assertion have already been investigated and proven less effective or ineffective. In summary, every recommendation must be judged on its own merits; hypothesized mechanisms of effect do not determine whether an outcome occurred, and the best choice is the recommendation that the best evidence (research) has demonstrated to result in the best outcomes (expected value/relative efficacy). The goal is not to find flawless recommendations; such an option may not even exist. The goal is to find the best possible recommendation.
"Best" is a Measurable Quantity
Effectiveness is a measurable quantity. The primary hypothesis asserted by this article is that there is an attainable best resistance training model that our industry should strive to achieve. This "best" model may be defined objectively and mathematically, rather than subjectively, using reliable, objective outcome measures. Further, the best outcomes may be achieved by recommending the best ranges for the modifiable variables in a resistance training program. It is important to note that "best outcomes" refers to the best average outcomes across all individuals (Axiom 1).
The most influential contributors to average outcomes are likely the reliability and magnitude of a recommendation's effect. These can also be described as frequency (how often the recommendation is effective) and value (the size of the effect). Their product is known as the expected value. The formula is expressed as: Frequency (reliability) x Value (effect size) = Expected Value (effect on outcomes). Expected value, a concept commonly used in game theory and economics, helps address the challenge of comparing recommendations that differ in both reliability and magnitude (4). For example, without this formula, it would be difficult to objectively compare a repetition range that is highly reliable but produces small effects to a set strategy that produces large effects but is rarely effective. By using the product of these two variables, we can more clearly evaluate recommendations and prioritize those most likely to improve average outcomes.
Practitioners and exercisers often fall prey to the availability heuristic, giving undue weight to memorable outcomes, such as dramatic successes or failures, while ignoring the frequency of those outcomes. This bias is particularly prevalent among professionals who are committed to a specific method or branded system. For example, a training method that produces a remarkable result for one client may only achieve this outcome 1 out of 10 times. Despite the strong anecdotal impression, its average effectiveness may be lower than that of more reliable alternatives. An example of this may be a trainer's implementation of lengthened partials to aid in breaking through a plateau in their own hypertrophy or strength development. However, comparative research on range of motion (ROM), hypertrophy, and strength suggests that lengthened partials are less effective than full ROM repetitions (see Range of Motion and Hypertrophy: Systematic Review). This does not mean that lengthened partials should be excluded from all programs. However, this result likely implies that any improvement following the addition of lengthened partials was due to other factors (for example, extending sets beyond full ROM failure, which functions like a drop set), and that lengthened partials alone are not a reliable strategy for improving outcomes.
Outcomes over Mechanisms: Why Can't We Base Recommendations on the Intended Effect?
Outcomes depend on the variables we can modify, rather than on our understanding of how those variables affect outcomes (mechanism of effect). The Bradford Hill Criteria of Causality, published in 1965, documented this interesting logical distinction. The Bradford Hill Criteria outline several factors that strengthen a hypothesis of a causal relationship (5, 6). One criterion is a supportable causal hypothesis; however, Bradford Hill also notes that it is not necessary to know how a variable affects an outcome to know that it does, in fact, affect an outcome. That is, knowing how something works is not required to know that something works.
Some everyday examples of this logic include making a phone call without understanding telecommunications, reducing a headache with aspirin without knowing its pharmacodynamics, or reaping the benefits of resistance training without understanding exercise physiology. In program design, adjusting a modifiable variable may improve an outcome regardless of whether the practitioner understands or correctly identifies the underlying mechanism of effect. The mechanism is a hypothesis about why a change occurs, while modification of variables focuses on what reliably produces better outcomes.
Further, an exerciser can benefit from the effects of a recommendation even if both the exerciser and their coach believe in an inaccurate causal explanation. For example, periodically changing exercise selection following a period of improvement for a particular exercise may improve outcomes even if both parties believe the benefit is due to "muscle confusion." The outcome does not depend on the accuracy of the explanation; the outcome depends on the variable that was modified.
It is important to understand that hypotheses do not affect outcomes unless they affect behavior, specifically, the setting of modifiable variables. Therefore, we must base recommendations on measured outcomes, not on intent or proposed mechanisms. This is not to say that mechanisms of effect are irrelevant; their value lies in generating new hypotheses or helping us identify additional variables worth modifying. In summary, outcomes improve when we modify variables in a way that increases effectiveness. Hypothesized mechanisms are useful only if they lead us to discover or refine the way we modify those variables.
A Program's Name or Branding Does Not Affect its Outcomes
A program's brand name is not a variable that affects outcomes. Every program, branded or not, is a collection of recommendations for setting the same modifiable variables (the central premise of this article), and two programs with the same settings are expected to produce the same outcomes, regardless of which name is most popular (Axiom 3). While it may be argued that one brand's coaching, community, or presentation improves adherence, this argument pertains to whether the program is performed, not to the brand itself as a factor that modifies outcomes. This insight has a practical implication: when a branded program produces good outcomes, the credit belongs to its recommendations for modifiable variables, and those recommendations can be identified, compared to the research, and improved. The best program is determined by the relative effectiveness of its recommendations, not by its marketing, its certification, or its celebrity endorsement.
Chance of Multiple Best Solutions
The previous article asked whether more than one "best approach" could exist, and concluded that it is exceedingly unlikely. In resistance training, the same question has a different answer, and the difference is instructive. Multiple best programs are not an unlikely exception in this model; multiple best programs are a defining feature of it (Axiom 6). Because outcomes are similar for a recommended range (e.g., 8-12 reps/set) for most modifiable variables, many programs tie for best, and the best region (V*) is precisely the set of all tied programs. The rehabilitation model expected a single winning combination of interventions; this model expects a single region of best programs that include the recommended range for each modifiable variable.
However, ties within the region and ties between regions are different claims, and the theorem above proves only the first. Ties within the region are guaranteed; that is what a range represents. A second, separate best region, a combination of settings far outside the current recommendations that produces equally best outcomes, remains exceedingly unlikely, for the same reason given in the previous article. Two distant programs would have to produce statistically indistinguishable and maximally effective outcomes across every variable simultaneously, and as the number of variables increases, the probability of this coincidence decreases dramatically. This is analogous to a multi-player poker game and the unlikely event that two players hold different, but equally winning, hands. In summary: many programs tie for best, but the tie exists inside one region, and the existence of a second, distant region of equal height is improbable enough to be discounted until evidence demonstrates otherwise (Axiom 12).
Threshold effects (points at which further increases in a variable no longer improve outcomes) were noted in the previous article, implying that certain thresholds of improvement may result in multiple solutions that yield the best possible outcomes. However, as mentioned above, the recommended range for each modifiable variable includes the range of values that result in the same outcomes. The difference in this model is that rather than a single best intervention combination, you end up with a best region, V* (the set of best possible programs on the outcome surface). For example, research on sets per muscle group demonstrates a plateau directly: improvements increase up to approximately 3 sets per muscle group per session, increase less with the addition of a 4th and 5th set, improve little beyond that range, and may decline at 6 or more sets (Axiom 5). Similar plateaus appear across the other modifiable variables; for example, ranges of repetitions, rest intervals, and tempos produce similar outcomes for a given goal. A range represents a plateau of recommendations and outcomes (Axiom 6); however, the width of each plateau is measured by comparative research, not assumed.
Ties inside the best region also have a practical benefit: preference can operate inside the region without sacrificing outcomes. If several programs within V* are expected to produce equivalent outcomes, an individual may select among them based on enjoyment, equipment, schedule, or variety, and the expected outcome does not change. This is the resistance training version of the previous article's conclusion on patient preference: preference may shape the final selection, but preference does not change the optimization used to identify the best available options. The model constrains the choices to the best region; the individual chooses freely within it.
Program Design is Not Zero-Sum, but Recovery is a Limited Resource
Resistance training program design is not a zero-sum problem. A zero-sum game is a scenario in which one choice is selected at the expense of another; every gain in one place requires a loss somewhere else. Intervention selection in the previous article was zero-sum because a session is only so long: every technique chosen consumed time, and every technique chosen was a technique excluded. Program design does not work this way. Modifiable variables do not compete for slots in a program, because every variable is always set in every program (the central premise of this article). Choosing a repetition range does not exclude choosing a tempo; both must be chosen, along with every other variable, in every program ever written. There is nothing to leave out, so there is no trade of one variable for another.
However, one resource in resistance training does create a limit, acting like a "budget": recovery. Recoverable capacity is the total training stress the body can recover from while continuing to adapt (Axiom 5), and the volume-related variables all draw from it. Sets, training frequency, and proximity to failure each add training stress, and the stress they add accumulates across the program. This is why the interactions between the volume-related variables behave like trade-offs even though the problem is not zero-sum overall (Axiom 4): the number of sets should be reduced when drop sets are introduced, and volume per session should be considered when frequency increases, because these variables share the recovery budget. In short, the variables do not compete with each other for a place in the program, but the volume-related variables do compete with each other for the same recovery capacity.
The value of spending recovery is governed by diminishing marginal utility. Diminishing marginal utility is the principle that each additional unit of a resource produces a smaller benefit than the unit before it. The research on sets per muscle group demonstrates this principle directly: improvements increase up to approximately 3 sets per muscle group per session, increase less with the addition of a 4th and 5th set, improve little beyond that range, and may decline at 6 or more sets (Axiom 5). Volume beyond the point of meaningful benefit is sometimes referred to as "junk volume": sets that cost recovery but purchase little or no additional improvement. The model's volume recommendations are the ranges in which each set purchases meaningful improvement, based on all available comparative research (Acute Variables: Sets per Muscle Group ).
The difference between a time budget and a recovery budget explains several practical realities of resistance training. A time budget ends when the session ends; a recovery budget spans the hours and days between sessions. This is why training frequency is a modifiable variable at all: sessions draw from a shared capacity that refills over time, so the distribution of volume across the week matters, not only the volume within a session. This is also why the best program is not the most program. When progress stalls, a common instinct is to add sets, add days, and move closer to failure. If the current program is already near recoverable capacity, each of these additions spends recovery that is not available, and outcomes may decline rather than improve (Axiom 5). Under a time constraint, doing more was impossible; under a recovery constraint, doing more is possible, and that is exactly what makes it a mistake.
In the formal model, this entire section is one line: R(v) ≤ Rₘₐₓ. The recovery cost of the program must not exceed recoverable capacity. The rest of the model is free to maximize expected outcomes in any direction; this constraint is the boundary it must respect while doing so.
The Problem with Current Practice: Default Settings and Practitioner Preference
If recommendations are not deliberately set based on comparative research, then the settings of most modifiable variables in most programs are, at least in part, accidental. Recall the central premise of this article: every program sets every modifiable variable, whether or not the program's author made that choice deliberately. A program built around a favorite repetition range and a preferred split still sets tempo, proximity to failure, rest intervals, exercise order, range of motion, and periodization; whatever the author did not consciously choose was set by habit, convention, equipment, or accident. The previous article compared best possible outcomes resulting from unprioritized intervention selection to expecting rolled dice to land in order from highest to lowest. The resistance training version is a machine with a panel of dials, one for each modifiable variable: a program author may deliberately set two or three dials, but the remaining dials do not wait patiently at "neutral." The remaining dials are already set, at whatever positions habit and convention left them, and there is no reason to expect accidental settings to land inside the best region.
Importantly, this problem is often not due to a lack of research. Comparative research exists for nearly every modifiable variable in this model, and for most variables the research base is substantial. A program author who has never considered optimizing rest between sets is not lacking the ability to make an evidence-based decision; the author is missing the application of evidence that already exists (Axiom 11). This mirrors a conclusion from the previous article: far more comparative research exists than is currently applied in practice.
Practitioner preference compounds the problem of accidental settings. Many professionals set variables based on allegiance to a training style or "school of thought," rather than on comparative research. For example, some coaches program exclusively heavy loads and low repetitions for every client and every goal, citing a strength-focused philosophy, despite comparative research suggesting that other repetition ranges produce better outcomes for hypertrophy and endurance (Acute Variables: Training Load (Weight and Resistance) ). Some coaches take every set to failure, citing an intensity-focused philosophy, despite comparative research suggesting that reps-in-reserve may produce better outcomes for increasing power and for athletes training at high frequencies and volumes (Acute Variables: Sets to Failure ). It is, of course, possible that a preferred setting is the best setting for a particular individual. However, preference is only harmless when it lands inside the best region. A deliberately set program constrains every variable to the research-supported range first, and then applies preference within those ranges, where preference costs nothing (as discussed above in Chance of Multiple Best Solutions).
The solution to accidental settings is not memorizing a single perfect program; the solution is the model: the Evidence-based Resistance Training Model (EBRTM) . A model that states the research-supported range for every modifiable variable, for each goal, converts program design from remembering the two or three variables an author cares about into checking every modifiable variable for its optimal range. The author's attention is no longer the limit on program quality; the model includes the variables the author would have forgotten, and the interactions between variables that the author would not have known to look for (Axiom 4).
Individualization Within the Model: Objective Outcome Measures
Every recommendation in this model is a population-level claim: the range most likely to get the best result for the greatest number of people (Axiom 1). Individuals vary, and some individuals will respond best to settings that differ from the population recommendation (different responders). This raises a reasonable question: if individuals vary, why start with the population recommendation at all? The answer is the same thought experiment presented in the "Best Physical Therapy Approach ." Imagine a repetition range that produces the best outcomes for 70% of individuals, and an alternative range that produces the best outcomes for the other 30%. Without additional information, no rational coach would start a client on the 30% option, and no client would choose it. There is no reason to assume a given individual is a different responder from the outset, and if the recommended range proves less effective, the alternative remains available. Starting inside the best region ensures that the greatest number of individuals receive the most effective settings in the fewest number of attempts.
Identifying a different responder requires measurement, and the measurement must be objective. In the previous article, this measurement was determined by reliable objective outcome measures (e.g., overhead squat assessment , goniometry , or validated functional assessment questionnaires); in resistance training, reliable, objective outcome measures include load lifted/height achieved, repetitions completed for a set at a given load, and performance on subsequent sets (total volume). Subjective impressions are poor substitutes. Perceived effort, soreness (which typically decreases rapidly as the body adapts to a new routine), and how a session felt are poor substitutes for measured performance; a program can feel productive while producing little benefit, or feel unremarkable while producing steady improvement. Coaches and exercisers who adjust programs based on feel are vulnerable to the same availability heuristic discussed earlier: memorable sessions outweigh measured trends.
A measurement should only be collected if changes in the measurement change recommendations. The previous article proposed a two-question test to determine whether an assessment was relevant, and the questions apply directly to program design: "What will I change if this measurement improves?" and "What will I change if this measurement does not improve?" If the answers to those two questions are the same, the measurement does not affect recommendations, and tracking the measurement results in record-keeping without adding value (it is not relevant). Applied to resistance training, tracking improvements in load lifted or repetitions per set passes these tests (stalled progression or a decrease in performance is the trigger for a change in recommendations), while mood, soreness, and even daily fluctuations in weight and anthropometric measurements may not be sufficient to trigger a change in recommendations.
Within the EBRTM, this measurement-and-adjustment process already operates continuously on one variable: autoregulated load adjustment, in which performance in each session is used to adjust the load recommendation up or down (Axiom 12). The same logic extends to the other modifiable variables: begin every variable inside the recommended range, measure outcomes objectively, and adjust an individual's recommended modifiable variable range when the measurements, not subjective impressions, indicate that the individual is a "different responder." Note the order of operations: the model supplies the starting settings, relevant assessment demonstrates efficacy, and changes in recommendations are based on the outcomes of relevant assessments. Measurement without the model starts individuals at accidental settings and adjusts from a bad position; conversely, the model without measurement holds every individual at the population recommendation and never detects "different responders." The best combination starts every individual at the settings most likely to be best and moves the different responders to their best settings in the fewest number of adjustments.

Section 2: The Best Recommendation Should Be Determined by Comparative Research Whenever Possible
To determine the best recommendation for each modifiable variable with the greatest accuracy, it is necessary to base relative effectiveness on the most reliable data available. Unfortunately, there is rarely enough data to precisely determine the reliability and average effect size (i.e., expected value) of every possible setting of every variable. However, it may be assumed that average outcomes (expected value) are the product of reliability and effect size. The most accurate source for estimating these expected values remains peer-reviewed and published research (referred to throughout this article as "research"). Research is the best tool available for minimizing bias and error, because it applies the largest number of controls (e.g., statistical analysis, blinding, peer review, independent replication) against the biases that practitioners cannot fully overcome in practice (e.g., confirmation bias, availability heuristic, anchoring). Further, because our goal is to recommend the range that produces the best outcomes relative to all other ranges, the research used must be comparative (Axiom 9).
How Research is Interpreted
The approach to interpreting research used to build this model has been published in detail, and only a summary is provided here (Using Research for Better Practice: A Decision Theory and Information Theory Approach ). In short: all relevant peer-reviewed research is included by default ("total evidence"), because dismissing studies based on flawed "levels of evidence" hierarchies shrinks the data set and increases the risk of error rather than reducing it (Levels of Evidence are Flawed ). Studies are then sorted and labeled into comparable subcategories (e.g., by population, training experience, outcome measure, duration, etc.), because apparent contradictions in research are usually category errors: studies that seem to disagree are often answering different questions. Within each subcategory, the direction of effect is extracted from each study (favors A, favors B, or no significant difference), and the trend is determined by vote counting (Axiom 10). Meta-analysis is not treated as a higher form of evidence; it is a different tool, reserved for cases where the direction of effect is unclear, or where a more precise estimate of magnitude could plausibly change the recommendation. The inappropriate elevation of meta-analyses over clear directional trends has contributed to a nihilistic view in exercise science ("nothing matters, everything works"), which is the research-interpretation version of a fallacy addressed earlier in this article; that is, demonstrating that many settings produce some improvement does not demonstrate that all settings produce the best improvement (Meta-Analysis Problems: A Crisis of Misuse and Misinterpretation ) (see Unsupported Default Position Fallacy).
The following is the vote-counting rubric used by the Brookbush Institute for all systematic reviews and educational content:
- Comparison Rubric (The goal is to use available research to determine the most likely trend.)
- A is better than B in all studies → Choose A
- A is better than B in most studies, and additional studies show similar results between A and B → Choose A
- A is better than B in some studies, and most studies show similar results between A and B → Choose A (with reservations)
- Some studies show A is better, some show similar results, and some show B is better → Results are likely similar (unless there is a clear moderator variable such as age, sex, or training experience that explains the divergence)
- A and B show similar results in the gross majority of studies → Results are likely similar.
- Some studies favor A, others favor B → Unless the number of studies overwhelmingly supports one side, results are likely similar.
Note how well this rubric fits the resistance training problem. When the rubric concludes "results are likely similar" across a span of settings, that conclusion is not a failure to find an answer; that conclusion is the answer. A span of settings with similar results is a plateau, and a plateau is what the model represents as a recommended range (Axiom 6). Vote counting does not only identify winners; vote counting measures the width of the best region. Sorting and labeling perform the same double duty: they resolve apparent contradictions, and they reveal the interactions between variables (Axiom 4). For example, when comparative studies of periodized and non-periodized programs are treated as one pile, the research looks "mixed"; when the same studies are sorted by training experience, a clear trend emerges, with only 1 of 13 studies favoring periodization for novice participants, and 9 of 17 studies favoring periodization for experienced participants (Periodization Training: Who needs it? ). One sorting step converted "periodization is controversial" into a recommendation with conditions.
Research Must Be Comparative
For the purposes of recommending the best range for a modifiable variable, the research must be comparative, directly evaluating the relative efficacy of at least two settings of that variable (Axiom 9). Comparative research may include both experimental and observational designs, with or without control groups. While randomized controlled trials (RCTs) benefit from increased internal validity, they are not strictly necessary for our purposes; for the purposes of comparing the relative efficacy of settings, settings must be compared. Note that non-comparative studies, including studies that compare a program to a randomized control group (no training or baseline), are less useful; demonstrating that training is better than not training does not demonstrate which settings are best. Furthermore, separate studies should not be directly compared to one another, due to the indeterminate number of confounding variables that may influence outcomes across studies. For example, if one study demonstrates a 5% strength improvement from 3 sets, and a different study demonstrates an 8% strength improvement from 5 sets, the difference cannot be attributed to sets; the studies may differ in population, training experience, exercise selection, duration, measurement method, and any number of unreported variables. Only a study that tested 3 sets against 5 sets within the same design can attribute the difference to sets.
Exact Values Are Not Required: Rankings Are Sufficient
A reasonable objection to the formal model is that the expected value of a program, E(v), can never be calculated exactly, because research rarely provides precise values for reliability and magnitude. The objection is correct about the numbers and wrong about the requirement. Locating the best region does not require numeric expected values; locating the best region requires knowing, for each modifiable variable, which settings produce better outcomes, which produce worse outcomes, and which produce equivalent outcomes. These are ordinal comparisons (rankings), and ordinal comparisons are exactly what comparative research provides and what vote counting extracts (Axioms 9 and 10). When the rubric concludes "Choose A," the setting A is ranked above the setting B. When the rubric concludes "results are likely similar," the settings are tied, and a span of tied settings is a recommended range (Axiom 6). Ranking every setting of every variable and identifying the interactions between variables (Axiom 4) bounds the best region without a single exact number.
This is a structural difference from the previous article worth noting. In the rehabilitation model, techniques competed for limited session time, so techniques had to be prioritized against each other across the whole pile, and prioritization across many items benefits from a scoring system that approximates cardinal values (a rank-weighted score). In this model, settings of one variable are only compared to other settings of the same variable, so a per-variable ranking is sufficient, and no scoring system is required. The one place magnitude matters more than direction is the recovery constraint: how much improvement each additional set purchases determines where volume stops being worth its recovery cost, which is why the sets-per-muscle-group research is discussed in terms of the size of improvements, and why meta-analysis is reserved for exactly these cases, where a representative magnitude could change the recommendation (Axiom 10). The nested selection subproblem is the other exception: choosing exercises from the pile of available exercises is a selection problem (see Same Family, Different Problem), and exercise selection can be prioritized by relative efficacy using the previous article's methods.
When Additional Information is Needed
Although research should serve as the primary source for determining the best recommendation, not all settings have been investigated in peer-reviewed and published studies, nor have all settings been directly compared to other settings. Research may permit some indirect comparisons; for example, research may suggest that 3 sets produce better outcomes than 1 set, and that 1 set produces better outcomes than untrained controls, leading to the assumption that 3 sets are superior to no training. While this inference is reasonable, indirect comparisons must be used cautiously, because settings may have varying levels of efficacy for different populations, training levels, or goals, and any conclusions drawn from them should be treated as provisional estimates until more direct comparisons are available (Axiom 12).
Further, an absence of evidence is not evidence of absence. A lack of research does not imply that a setting is ineffective; it only implies that its relative expected value cannot be estimated with research. When research is unavailable, in-practice comparisons become essential, and these comparisons should not be based on intuition or gut-level impressions. Instead, in-practice comparisons should be based on reliable, objective outcome measures, tracked over time, ideally testing settings within the same individual and repeated across multiple individuals before drawing conclusions (as discussed in Individualization Within the Model, and expanded in the experimentation section below).
There Is More Research Available Than Is Currently Utilized
One critique of this framework is that, if a single best model exists, the current body of research is insufficient to identify it. While substantial gaps remain in the literature, far more research exists than has been previously applied in practice. The Brookbush Institute was founded with the intent of developing the first comprehensively evidence-based education platform, with every course built from a systematic review of all relevant peer-reviewed research on that topic. The Evidence-based Resistance Training Model (EBRTM) was constructed this way: a systematic review for each modifiable variable, built from all of the relevant comparative research available at the time of the review. Without exception, every variable reviewed revealed dozens, and sometimes hundreds, of studies that were overlooked in previously published reviews and educational materials. These studies address nuanced questions, contribute critical details, and clarify apparent contradictions; in at least one case, the overlooked details revealed a variable that traditional acute variable lists omitted entirely (proximity to failure). The persistent critique that there is "not enough research" overlooks the substantial gains in accuracy that could be achieved by fully leveraging the existing research (Axiom 11). Even if this process falls short of identifying the global optimum, it represents the closest current approximation, and the refinement process (Axioms 9-12) continues from there. The point is not perfection, but optimization; continual refinement through the comprehensive application of current evidence.

Section 3: Refining and Confirming the Model
Experimentation is Necessary to Confirm the Optimum is Global
A local maximum is a solution that is the best of everything tried so far; a global maximum is the best solution possible. The current best region is built from comparative research, and comparative research can only test the settings researchers chose to test (Axiom 12). If a better setting exists outside every tested range, no amount of reviewing can find it, because the evidence does not exist to be reviewed. The current best region may therefore be a local maximum: the highest region on the mapped portion of the outcome surface, with unmapped territory remaining. Discovering whether better outcomes are possible requires deliberately testing settings outside the current best region and measuring the results. Only exploration can confirm that the current best region is the global maximum, or replace it with a better one.
A note of caution before describing the method: experimentation should begin only after the model has been learned and used. While it is technically possible for a novice to stumble upon a more effective setting, such discoveries are unlikely in a field with a large body of comparative research. It is far more likely that individuals with limited knowledge of the research will unknowingly retest settings already demonstrated to be inferior, or abandon promising settings because no objective outcomes were tracked. This is another consequence of completeness (Axiom 11): knowing what the research has already tested is what separates exploration from repetition. An experiment that overlaps with settings already represented in the research does not add new territory to the map; an experiment chosen with knowledge of the research begins where the map ends (Axiom 12).
The method is to change one variable at a time, and to measure the change objectively. An experiment in resistance training is structured as follows. First, set every modifiable variable inside its recommended range, and train long enough to establish a measured baseline of progression (for example, load and repetitions tracked across several weeks). Second, move one variable, and only one variable, to a setting outside its recommended range, and hold every other variable constant. Third, continue the same objective measurements for a full training block (for example, 4-8 weeks), long enough for a trend to emerge. Fourth, compare the measured progression during the experiment to the measured baseline. Autoregulated progressions make this comparison practical: because load already adjusts to performance session by session, the rate of progression is continuously recorded, and a change in that rate is the experiment's result. Changing one variable at a time is what makes the result interpretable; if two variables change and progression improves, the experiment cannot determine which variable deserves the credit.
This structure limits the cost of being wrong. The previous article compared structured experimentation to a strategy from finance called the barbell strategy, described by Nassim Nicholas Taleb: the majority of a portfolio is held in conservative, low-risk investments, while a small portion is allocated to speculative, high-risk opportunities (7). The resistance training version holds every variable but one inside the research-supported ranges (the conservative majority) while one variable explores outside its range (the small speculative allocation). Even if the experimental setting is ineffective, a program with every other variable inside the best region remains a good program, and the cost of the experiment is limited to the difference on a single variable for a single training block. The potential reward is a discovery the research has not yet made.
The results of an experiment update the model at two levels, and the two levels should not be confused. At the individual level, an experiment that produces better measured outcomes identifies a different responder, and that individual's settings should be adjusted (as discussed in Individualization Within the Model). A single individual's result does not change the population recommendation, for the same reason one testimonial does not validate a program (Axiom 1). At the model level, recommendations change when comparative research changes: when new studies test new settings, the reviews are updated, and the best region moves if the evidence moves it (Axioms 9-11). Practitioner experimentation contributes to the model level by generating hypotheses worth testing formally, and the settings that repeatedly outperform expectations across many individuals are exactly the settings researchers should test next. This is the full version of the process that autoregulated load adjustment performs continuously on a single variable: measure, compare, and update, with the exploration now aimed at the unmapped portions of the outcome surface.
Same Family, Different Problem: Comparing the Two Sets of Axioms
The axioms section promised a comparison of this article's axioms with the axioms of the previous article, and the comparison is worth making explicit, because the overlaps and the differences prove different claims. The overlaps demonstrate that the two articles are one method: a shared definition of "best," shared rules of evidence, and a shared recognition that each field is an optimization problem. The differences demonstrate that the two articles solve different optimization problems. What changed between the articles is exactly what makes them different optimization problems, and nothing else. If the two axiom sets were identical, resistance training would just be the physical rehabilitation proof with different nouns. If the two sets shared nothing, the claim of being a common method would be false.
The previous article originally defined intervention selection with ten axioms. Summarized briefly for readers of this article only: outcomes are probabilistic; choices are relative; "best" is a measurable quantity (expected value); selection is zero-sum, because session time is limited; maximizing the sum of expected values maximizes outcomes; expected values are initialized from comparative research; deliberate experimentation is necessary to avoid local maxima; diminishing marginal utility governs additional interventions; the system converges toward optimal intervention sets; and assessment enables subgroup-specific optimization.
Six commitments are shared between the two articles, and together they form the method of this series. Outcomes are probabilistic (Axiom 1 in both articles). "Best" is highest expected value, which also makes every evaluation relative rather than binary (this article's Axiom 2, covering the previous article's second and third axioms). Values are initialized from comparative research (Axiom 9). Comparative findings are synthesized by vote counting within sorted subgroups (Axiom 10). Completeness raises the probability of finding the true optimum (Axiom 11). Deliberate experimentation is required to confirm the optimum is global (Axiom 12 in this article; Axiom 10 in the previous article). Note that the original version of the previous article applied vote counting and the completeness argument in its methods discussion rather than stating them as axioms; the updated version now states both as axioms, aligning the two sets. These six commitments define "best" and the rules of evidence, and they transfer to any field in which options can be compared on outcomes.
The most important shared commitment may be the one that is not an axiom at all: the recognition that both fields are optimization problems. An optimization problem has three fundamental components: decision/modifiable variables (the things you can control), an objective function (the measure of success to be maximized), and constraints (the limits on allowable choices). Solving an optimization problem involves finding the best possible solution from all available choices. Recognizing these three components in each field is what allows centuries of optimization mathematics to be borrowed rather than reinvented. Comparing the two articles by these three components demonstrates why they are different optimization problems.
First, the decision/modifiable variables differ. In rehabilitation, the decision variables were selections; that is, which techniques, from the pile of all available techniques, would be included in the session (a yes or no for every technique). In resistance training, the decision variables are settings: a recommended range for each modifiable variable, where every variable must be set in every program (Axiom 3). This difference produces two further structural axioms for optimizing resistance training. Because the variables are set simultaneously rather than chosen against one another, the variables interact, and the interactions are part of the model (Axiom 4). Because outcomes are similar across a span of settings, the best setting of each variable is a range, not a point (Axiom 6).
Second, the objective function differs. In rehabilitation, the objective was a sum: the total expected value of the session is the expected values of the selected techniques added together, and the best session is the combination with the highest total. In resistance training, the objective is an "outcome surface": the expected outcome is a function of all of the settings together, and the best program is the best range for each setting, resulting in the highest region on that "outcome surface" (Axiom 7). Further, resistance training has four objective functions, one for each goal (endurance, hypertrophy, strength, and power), and the goal is chosen by the individual before optimization begins (Axiom 8); rehabilitation is optimized toward a single objective, recovery.
Third, the constraint differs. In rehabilitation, the constraint was time: a session is only so long, so every technique chosen was a technique excluded, and selection was zero-sum. In resistance training, the constraint is recovery: the body's capacity to adapt to training stress is limited, so the volume-related variables share a recovery budget even though nothing is ever excluded from a program (Axiom 5). Two of the previous article's axioms were absorbed by this replacement rather than discarded: the zero-sum axiom became the recovery constraint, and diminishing marginal utility became a property of that constraint, because each additional set results in less improvement than the set before it.
The remaining differences follow the same pattern of reassignment rather than removal. Assessment-driven subgrouping became two structures in this article: the goal, which is the largest and most consequential subgroup in resistance training and received its own axiom (Axiom 8), and objective outcome measures, which identify different responders (the section Individualization Within the Model). Last, the previous article's convergence axiom is not an axiom in this article at all. The existence of a single best model, and the convergence of refinement toward it, is derived from the other axioms as a theorem. This is a structural strengthening: a claim assumed in the original version of the previous article is now proved in both versions; the updated version of the previous article derives its conclusion in the same way.
A note worthy of additional study: the two problem types nest inside each other. A savvy reader may observe that each article's problem contains a small version of the other article's problem. In resistance training, exercise selection is a modifiable variable (Axiom 3), and setting that variable is itself an intervention selection, like the rehabilitation problem of choosing exercises from the pile of all available exercises. In rehabilitation, once a technique is selected, that technique must be dosed, and dosing is itself is a "training recommendation": a therapeutic exercise still requires a load, repetitions, sets, and tempo. This nesting does not weaken the classification of either problem; the classification describes the dominant structure of each field's decision problem, which determines the mathematics of the model as a whole. However, the nesting does make a useful prediction: wherever a subproblem of the other type appears, the other article's methods apply to that subproblem. Exercise selection within this model can be prioritized by relative efficacy, exactly as interventions were prioritized in the previous article; the dosage of a rehabilitation exercise can be set by recommended ranges, exactly as the modifiable variables are set in this article. The two articles are not only two examples of one method; each is also the instruction manual for the subproblems inside the other.
In summary, the two articles share their definition of "best," their rules of evidence, and their recognition of an optimization problem; they differ in every component that describes the shape of the problem: the decision variables, the objective function, and the constraint. That is precisely the pattern expected if the method is sound: the method holds constant across fields, and the problem structure is dictated by the field. The pattern also implies a recipe that extends beyond this series. Any field in which options can be compared on reliable, objective outcomes can be given the same treatment: state the shared commitments, identify the three components of the field's optimization problem, add the axioms the structure requires, and derive whether a best approach exists.
The most obvious candidates for this treatment are the other fields of medicine; for example, primary care, internal medicine, and their specialties. The components already exist in medicine, in isolation. Clinical decision analysis has applied expected value to individual decisions since the 1970s; for example, the threshold model of Pauker and Kassirer uses expected utility to determine whether to treat, test, or withhold treatment for a single condition (8). Axiomatic models of individual treatment choice have also been published (9), and the closest existing program, the threshold and decision-theory work of Djulbegovic and Hozo, shares this series' commitment to expected value at the point of care; however, that program optimizes one decision at a time, and it does not derive that a best approach exists for a field (10). Constrained optimization is an established method in health economics and policy, where it is used to allocate limited budgets across screening programs and interventions (11). Sequencing models have been used to optimize the timing of treatments for single diseases (12), and intervention-optimization methods use structured experiments to improve multicomponent interventions, although their authors state that an "optimized" intervention is not best in an absolute or ideal sense (13). Comparative effectiveness research exists as a named discipline with its own literature. Recently, an axioms-plus-theorem structure has appeared within a specialty for the first time: a 2026 paper proposed six axioms for ordering diagnostic urgency in emergency medicine and proved that any urgency function satisfying those axioms takes a specific form (14). Note that this work addresses the order in which diagnoses should be evaluated, not the selection or specification of treatment, and it makes no claim that a best approach exists for the specialty. To our knowledge, no medical specialty has combined these components the way this series combines them: stating the axioms of the specialty's decision problem, identifying which type of optimization problem the specialty faces, initializing the model from a comprehensive synthesis of the comparative research, and deriving that a best approach exists at the level of everyday practice (and modeling that approach), rather than at the level of budgets and policy. Each specialty would require its own analysis, because each specialty may be a different type of optimization problem: some may be selection problems like rehabilitation, some may be specification problems like resistance training, and some may be sequencing problems (problems in which the order of decisions over time is the primary variable; a type this series has not yet required). If this framework is as generalizable as the first two applications suggest, the articles in this series are not two solutions to two problems; they are the first two worked examples of a method for finding the best approach in any outcome-comparable field.
Section 4: The Current Best Region - The Evidence-based Resistance Training Model (EBRTM)
Everything above establishes that a best region exists, how it is estimated, and how it is refined. This section presents the current estimate. The Evidence-based Resistance Training Model (EBRTM) was developed from the Brookbush Institute's systematic reviews of every modifiable acute variable that significantly influences resistance training outcomes, integrating findings from approximately 1,500 peer-reviewed studies. This is many times the number of studies used in the development of other models and acute variable tables. The variables reviewed are the decision variables of this article's formal model: load, repetition range, repetition tempo, proximity to failure, rest between sets, circuit training, sets per muscle group, training frequency, set strategies, periodization, range of motion, exercise order, and exercise selection. Each variable was reviewed using the method described in Section 2: all relevant comparative research was gathered, studies were sorted into comparable categories, effect directions were extracted, and recommendations were developed from the trends (Axioms 9-11).
The model's central finding is a discovery about the shape of the outcome surface. Recall that this model has four objective functions, one for each goal (Axiom 8), which means resistance training has four outcome surfaces defined over the same variables, and the best region of each surface could, in principle, sit anywhere. The research demonstrates that the four best regions overlap on most variables. Endurance, hypertrophy, strength, and power share the majority of their programming, and differ on only a handful of variables. In the vocabulary of this article: the four best regions coincide on most axes of the surface, and a program moves from one goal's best region to another's by adjusting a small set of variables, not by rebuilding the program.
Brief Summary: Evidence-based Resistance Training Model (EBRTM)
Novice Recommendations
The model also demonstrates a second structural finding: the best region for novice exercisers is one region, regardless of goal. During the first 6-12 weeks of training, the recommendations converge for all four goals, and changing a novice's acute variables to match an experienced lifter's goal-specific model does not produce better outcomes. In the vocabulary of this article: for the first 6-12 weeks, the four outcome surfaces are effectively the same surface.
- Load and repetition range: Moderate loads (8-12 RM / 70-80% of 1-RM), with frequent autoregulated load adjustment.
- Tempo: Controlled eccentric (2 or more seconds) with a maximum-velocity concentric.
- Proximity to failure: Repetitions to failure (or 1-2 repetitions-in-reserve for those also performing sport or skill work).
- Range of motion: The largest range achievable with good form and without pain.
- Rest: 2-3 minutes between sets for similar muscle groups; 30-60 seconds between exercises in a circuit.
- Sets per muscle group per session: 1-2 for upper-body muscle groups; 2-3 for lower-body muscle groups.
- Frequency and split: 3 sessions in 2 weeks, progressing to 2 and then 3 sessions per week; full-body routines.
- Periodization and set strategies: None.
- Progression order: Form, then repetitions, then load, then sets, then exercise progression (stability).
- Exercise selection: Begin with relatively stable exercises.
- The one goal-specific exception: High-velocity exercise is introduced early for power goals, progressing by height or speed before adding eccentric load, stability, or external load.
Shared Variables
The shared variables are the recommendations that remain largely the same across endurance, hypertrophy, strength, and power for experienced exercisers. Each recommendation is a range, each range is a slice of the best region (Axiom 6), and each is supported by its own systematic review, cited in the model. A sample of the shared recommendations:
- Rest between sets: 2-3 minutes between sets for similar muscle groups; 30-60 seconds between exercises in a circuit.
- Range of motion: The largest range achievable with good form and without pain; a temporary reduction is acceptable if it allows a beneficial increase in load.
- Concentric intent: Every repetition is performed with the intent to move at maximum velocity, to maximize motor unit recruitment and force production.
- Exercise order: The most important, largest-muscle, and multi-joint exercises are performed first; resistance training before aerobic exercise, unless aerobic performance is the priority.
- Sets per muscle group per session: 2-5 sets, progressed over time.
- Training frequency: 1.5-3 sessions per muscle group per week, with 2-5 days between sessions.
- Training splits: Total-body routines at 1-2 sessions per week; an upper/lower split at 4 sessions per week.
- Periodization: Progresses through stages, from autoregulated load adjustment within one repetition range, to linear progression of intensity, to daily undulation of two repetition ranges, which becomes the ongoing template.
Goal-Dependent Variables
Goal-dependent variables are the small set of variables that differentiate the four goals: the primary variable progressed, load emphasis, proximity to failure, tempo, exercise selection priorities, and advanced set strategies. Summarized by goal:
- Endurance
- Primary progression: Reps to failure
- Load and Daily Undulation: Mostly light, some moderate, occasionally heavy
- Exercise selection: Challenge stability: light loads (6-10/10), moderate loads (4-6/10)
- Proximity to failure: To failure on most or all sets
- Tempo: 2+ : 0-2: MaxV
- Advanced set-strategy: Drop sets on moderate and heavy load days
- Hypertrophy
- Primary progression: Volume (load × reps)
- Load and Daily Undulation: Mostly moderate, some heavy, occasionally light
- Exercise selection: Moderate challenges to stability: moderate loads (4-6/10), heavy loads (1-3/10)
- Proximity to failure: To failure on most or all sets (unless also a high-volume athlete)
- Tempo: 2+ : 0-2: MaxV
- Advanced set-strategy: Drop sets on higher-volume days
- Strength
- Primary progression: Load
- Load and Daily Undulation: Mostly heavy · some moderate · occasionally light
- Exercise selection: Mostly stable exercise, focus on force production: heavy loads (1–3/10), moderate loads (4–6/10)
- Proximity to failure: To failure on most sets, or at least the last set; heavy days may use reps-in-reserve plus an extra set to preserve volume
- Tempo: Moderate to heavy loads - 2+: 0-2: MaxV; very heavy loads - lifter's preference (MaxV concentric)
- Advanced set-strategy: Inter-set rest intervals (≤20 s) may be used to perform more reps while maintaining force; accommodating resistance may be used to aid in breaking strength plateaus; drop set on the last set may be used to increase volume
- Power
- Primary progression: Velocity/height/distance
- Load and Daily Undulation: Mostly heavy loads and max power exercises, some moderate loads and power-stability exercises
- Exercise selection: Mostly stable exercises and high-velocity exercises: max power, stability power, heavy loads (1–3/10), moderate loads (4–6/10)
- Proximity to failure: 1–3 reps-in-reserve plus 1–2 extra sets to preserve volume; last-set-to-failure only when the session is not before practice/competition
- Tempo: Power: Explosive: 0: Explosive; Strength: lifter's preference
- Advanced set-strategy: Inter-set rest intervals (≤20 s) may be used to perform more reps while maintaining force and velocity; accommodating resistance during strength exercises to aid in breaking strength plateaus; drop sets on the final set of strength training (when not before practice or games).
The full recommendations for each goal, including tempo prescriptions, stability progressions, and routine construction, are presented in the model and its four goal-specific courses (Endurance Training: Evidence-based Model ; Hypertrophy Training: Evidence-based Model ; Strength Training: Evidence-based Model ; Power Training: Evidence-based Model ).
Two findings from the model's development deserve mention because they demonstrate the framework's claims in practice. First, the comprehensive review revealed a variable that traditional acute variable tables omitted entirely: proximity to failure. The reviews of load, repetitions, and set strategies repeatedly demonstrated that proximity to failure influenced outcomes more than the variables being studied, and it is now one of the six goal-dependent variables; this is the completeness argument (Axiom 11), producing a concrete discovery, because a variable cannot be set deliberately until it is identified. Second, the research demonstrates that several popular subgroups do not require separate models: gender does not influence the recommendations, and goals such as weight loss, general fitness, and longevity do not require their own programming, because the research suggests the choice of resistance training goal has little effect on those outcomes (Axiom 8).
Honesty about the model's boundaries is part of the method (Axiom 12). A small number of recommendations rest on trends observed across reviews rather than on a dedicated systematic review; for example, periodic changes in exercise selection (approximately once every 4-12 weeks) are recommended based on trends appearing across the acute variable reviews, and the model identifies this as a topic warranting a systematic review of its own. These recommendations are labeled provisional estimates, exactly as Section 2 requires. The model is the current best estimate of the best region, initialized from the most complete synthesis of comparative research yet performed for resistance training, and it will be refined as the research grows (Axioms 9-12); per the theorem, refinement converges toward the true best region rather than away from it.

Section 5: Conclusion
This article asserted that a single best resistance training model exists, and then did the work the assertion requires. "Best" was defined as highest expected value, an objectively measurable quantity (Axiom 2). Resistance training program design was identified as an optimization problem, and as a specific type of optimization problem: a specification problem, in which every modifiable variable must be set in every program, the variables interact, and the constraint is the body's capacity to recover (Axioms 3-8). Twelve axioms defined the problem completely, and the central claim was not asserted as an axiom; the existence of a single best model was derived from the axioms as a theorem. The best region was then estimated from the most complete synthesis of comparative research yet performed for resistance training, approximately 1,500 peer-reviewed studies reduced to recommended ranges for every modifiable variable, the Evidence-based Resistance Training Model (EBRTM) . Last, the article described how the estimate is refined: objective outcome measures identify different responders, structured experimentation tests settings beyond the current best region, and new comparative research updates the reviews (Axioms 9-12).
The structure of this argument determines the structure of any valid critique. Because the conclusion is derived from the axioms, rejecting the conclusion requires identifying which axiom is false; rejecting the conclusion alone is not available. Critiques of individual recommendations are different, and they are welcome. A recommendation in this model is a claim about the direction of effect across the comparative research, so a valid critique identifies research that was missed, a sorting error, or new studies that change the trend. Critiques of this form are not threats to the model; they are the model working, because the model is designed to be updated by exactly this kind of input (Axiom 12). What the model does not accept as critique is preference, tradition, mechanistic hypotheses that are not supported by outcomes, or the observation that some other program also "works" (see Critique is Not Binary, and the Unsupported Default Position Fallacy).
The practical implications are stated plainly. For the exerciser: every program ever performed set every variable in this model, and the only question is whether those settings were deliberate. Starting inside the recommended ranges results in the highest probability of the best outcomes; preference can operate freely within the ranges at no cost, and objective measurement, not subjective feelings, determines when settings should differ from recommendations. For the professional: program design is no longer a matter of allegiance to a training style; it is checking every modifiable variable against its research-supported range, and additionally, the model provides recommendations for the variables and relationships that attention alone may miss. For the industry: certifications, acute variable tables, and branded systems can now be evaluated against a model built from many times the research base of any previous recommendations, and the differences are verifiable, for each modifiable variable, against a comprehensive systematic review.
This article is the second worked example of a method. The previous article demonstrated that physical rehabilitation intervention selection is an optimization problem, known as a "selection problem," solved with the mathematics of constrained selection. This article demonstrated that resistance training program design is a "specification problem," solved with the mathematics of optimum conditions. The method that produced both is general: define "best" as expected value, identify the type of optimization problem, state the axioms the structure requires, derive whether a best approach exists, and initialize the model from a comprehensive synthesis of the comparative research. Fields throughout medicine may benefit from the same approach.
Note that the goal is not perfection, and the goal is not a flawless program; perfection and/or a flawless program may not exist. The goal is optimization: the best recommendations the current evidence can support, stated precisely enough to be checked, structured to be corrected, and improved continuously as the research grows. The theorem guarantees a best region exists. The model is the current best estimate of where it is. The method guarantees the estimate gets better.
Annotated Bibliography
- Dantzig, G. B. (1957). Discrete-variable extremum problems. Operations Research, 5(2), 266–288. This paper is the canonical formalization of the knapsack problem: selecting the subset of items that maximizes total value without exceeding a fixed capacity. Dantzig demonstrated that selection under a shared constraint is a distinct class of optimization problem with its own solution methods, and the 0/1 knapsack model described in this paper is the direct formal ancestor of the intervention-selection model proposed in the previous article in this series. The citation is included to establish that "choosing the best combination that fits" is not a metaphor but a problem class with roughly 70 years of published mathematics behind it.
- Box, G. E. P., & Wilson, K. B. (1951). On the experimental attainment of optimum conditions. Journal of the Royal Statistical Society: Series B (Methodological), 13(1), 1–45. This paper founded response surface methodology (RSM), and it asked the question this article adapts: how should the operating conditions of a process (temperature, pressure, concentration, time) be set to maximize yield, when every variable must be set in every batch, the variables interact, and the response changes smoothly enough that the optimum is a region approached by mapping a surface? Box and Wilson demonstrated that specification problems of this type are structurally different from selection problems and require different methods, including sequential experimentation toward the optimum. This citation is included to establish the lineage of the "outcome surface" framing; note that citing RSM is not performing RSM, in the same way that citing a meta-analysis is not performing one. The comparative research base plays the role of the accumulated experimental record, and practitioner-level experimentation (Axiom 12) is the fragment of Box and Wilson's sequential method that remains available when the published record, rather than the experimenter, chooses which settings get tested.
- Hubal, M. J., Gordish-Dressman, H., Thompson, P. D., Price, T. B., Hoffman, E. P., Angelopoulos, T. J., Gordon, P. M., Moyna, N. M., Pescatello, L. S., Visich, P. S., Zoeller, R. F., Seip, R. L., & Clarkson, P. M. (2005). Variability in muscle size and strength gain after unilateral resistance training. Medicine and Science in Sports and Exercise, 37(6), 964–972. This study measured the range of individual responses to an identical resistance training program in the largest cohort tested to date: 585 subjects (342 women, 243 men) performed 12 weeks of progressive unilateral elbow flexor training, with muscle cross-sectional area measured by MRI and strength measured by one-repetition maximum and maximal voluntary contraction. Changes in muscle size ranged from a 2% decrease to a 59% increase, and one-repetition-maximum strength gains ranged from 0% to 250%. This citation is included as the direct evidence for Axiom 1: an identical program produces a spread of outcomes across the individuals who perform it, from no measurable improvement to improvements several times larger than the group average, which is why every recommendation in this model is a probabilistic, population-level claim rather than a promise to an individual.
- von Neumann, J., & Morgenstern, O. (1944). Theory of Games and Economic Behavior. Princeton University Press. This book founded modern decision theory and game theory, and it formalized the principle this series borrows as its definition of "best": rational choice under uncertainty is the maximization of expected utility, derived in the book from a small set of axioms about rational preference. The citation is included to label what is borrowed as borrowed. Defining "best" as highest expected value is not an innovation of this series; the innovation is applying that definition to fields where "best" was still being decided by convention, professional preference, and allegiance to schools of thought. The book is also the structural precedent for this article's method: von Neumann and Morgenstern did not argue that expected-utility maximization is reasonable; they proved it follows from axioms, which is the same relationship this article establishes between its twelve axioms and its theorem.
- Hill, A. B. (1965). The environment and disease: Association or causation? Proceedings of the Royal Society of Medicine, 58(5), 295–300. This address introduced what are now known as the Bradford Hill criteria: a set of considerations (including strength, consistency, temporality, biological gradient, and plausibility) that strengthen a hypothesis of causation. The citation is included for one specific point within the criteria: Hill noted that a supportable causal mechanism strengthens a causal hypothesis, but that knowing how a variable affects an outcome is not necessary to know that it does affect the outcome. This is the published foundation for the Outcomes over Mechanisms section: recommendations are validated or refuted by measured outcomes, and mechanistic hypotheses matter only insofar as they change behavior or generate new variables worth testing.
- Fedak, K. M., Bernal, A., Capshaw, Z. A., & Gross, S. (2015). Applying the Bradford Hill criteria in the 21st century: How data integration has changed causal inference in molecular epidemiology. Emerging Themes in Epidemiology, 12, 14. https://doi.org/10.1186/s12982-015-0037-4 This paper reviews how the Bradford Hill criteria have been applied and reinterpreted in the 50 years since their publication, including the criterion of plausibility. The citation is included as the modern companion to Hill (1965), demonstrating that the distinction between evidence of an effect and explanation of an effect remains current in epidemiological practice, and providing readers a contemporary entry point into the causal-inference literature.
- Taleb, N. N. (2012). Antifragile: Things That Gain from Disorder. Random House. This book describes the barbell strategy: holding the majority of a portfolio in conservative, low-risk positions while allocating a small portion to speculative, high-risk opportunities, so that the downside of exploration is capped while the upside remains open. The citation is included because the experimentation method recommended in this article is a barbell strategy applied to program design: every modifiable variable but one is held inside the research-supported ranges (the conservative majority), while one variable explores a setting outside its range (the small speculative allocation), limiting the cost of a failed experiment to a single variable for a single training block.
- Pauker, S. G., & Kassirer, J. P. (1980). The threshold approach to clinical decision making. New England Journal of Medicine, 302(20), 1109–1117. https://doi.org/10.1056/NEJM198005153022003 This paper is a canonical application of expected value to medical practice: it derives testing and treatment thresholds from the probabilities and utilities of clinical outcomes, determining when a clinician should treat, test, or withhold treatment for a single condition. The citation is included in the discussion of extending this series' framework to medicine, as evidence that expected-value reasoning at the point of care has an established history. The threshold approach optimizes one decision at a time; it does not classify a specialty's everyday practice as a type of optimization problem, and it does not derive that a best approach exists for a field, which is the combination this series proposes.
- Karni, E. (2009). A theory of medical decision making under uncertainty. Journal of Risk and Uncertainty, 39(1), 1–16. https://doi.org/10.1007/s11166-009-9071-3 This paper presents an axiomatic model of medical decision making for the choice of treatment following a diagnosis, with the optimal treatment defined by expected utility. The citation is included as evidence that axiomatic treatment-choice models exist in the medical decision-making literature. Like the threshold approach, the model addresses an individual treatment decision; it does not axiomatize a specialty's practice as a whole, initialize a model from a comprehensive synthesis of comparative research, or derive the existence of a best approach for a field.
- Djulbegovic, B., & Hozo, I. (2023). Threshold Decision-Making in Clinical Medicine: With Practical Application to Hematology and Oncology (Cancer Treatment and Research, Vol. 189). Springer. This book consolidates the closest existing academic program to the framework proposed by this series: a sustained body of work applying expected utility theory, regret theory, and decision thresholds to point-of-care medical decisions, with explicit integration into evidence appraisal and guideline development. The citation is included to name the nearest neighbor accurately. The program shares this series' commitments to expected value and to decisions at the point of care; it differs in that it optimizes one decision at a time, does not classify a specialty's practice as a named type of optimization problem, and does not claim or derive that a single best approach exists for a field.
- Crown, W., Buyukkaramikli, N., Thokala, P., Morton, A., Sir, M. Y., Marshall, D. A., Tosh, J., Padula, W. V., Ijzerman, M. J., Wong, P. K., & Pasupathy, K. S. (2017). Constrained optimization methods in health services research—An introduction: Report 1 of the ISPOR Optimization Methods Emerging Good Practices Task Force. Value in Health, 20(3), 310–319. https://doi.org/10.1016/j.jval.2017.01.013 This task force report introduces constrained optimization methods (linear, integer, and nonlinear programming) to the health services research community, describing them as methods for identifying the best policy choice or clinical intervention given a goal and a set of constraints. The citation is included as evidence that typed, constrained optimization is an established method in health economics and policy. The report's applications operate at the level of budgets, screening programs, and resource allocation; the framework proposed by this series applies the same class of mathematics at the level of everyday practice.
- Denton, B. T., Kurt, M., Shah, N. D., Bryant, S. C., & Smith, S. A. (2009). Optimizing the start time of statin therapy for patients with diabetes. Medical Decision Making, 29(3), 351–367. This paper formulates the timing of statin initiation for patients with type 2 diabetes as a Markov decision process and solves for optimal treatment-initiation policies. The citation is included as a representative example of sequencing-type optimization applied to a single disease-management problem: a named optimization model, an expected-value objective, and a provably optimal policy, applied to one decision within one condition. It demonstrates both that the mathematics is ready for medicine and that existing applications address individual problems rather than a specialty's practice as a whole.
- Collins, L. M. (2018). Optimization of Behavioral, Biobehavioral, and Biomedical Interventions: The Multiphase Optimization Strategy (MOST). Springer. This book presents the Multiphase Optimization Strategy, an engineering-inspired framework for optimizing multicomponent interventions through structured factorial experimentation, defining optimization as identifying the intervention that provides the best expected outcome obtainable within constraints of efficiency, economy, and scalability. The citation is included for two reasons. First, MOST is the strongest published precedent for structured experimentation as a route to better interventions. Second, Collins explicitly states that "optimized" does not mean best in an absolute or ideal sense, which is the substantive counter-position to this article's central claim; the theorem in this article is the answer to that position, because it derives the existence of a best region from axioms that can be individually examined and challenged.
- Safari, P., Mohamed, A., & He, S. (2026). Axioms of diagnostic urgency: A characterization theorem. Entropy, 28(8), 853. This paper proposes six axioms for ordering diagnostic urgency in emergency medicine and proves a representation theorem: any urgency function satisfying the axioms can be represented by an additive score. The citation is included as an independent demonstration that the axioms-plus-theorem structure has begun to appear inside medical specialties. The paper addresses the ordering of diagnostic evaluation, not the selection or specification of treatment, and it makes no claim that a best approach exists for the specialty; it post-dates the first article in this series.
- International Council for Harmonisation. (2009). ICH Harmonised Tripartite Guideline Q8(R2): Pharmaceutical Development. This regulatory guideline formalized the concept of a "design space": a multidimensional region of process parameters within which any combination of settings has been demonstrated to produce acceptable product quality, such that movement within the region is not classified as a change requiring new approval. The guideline is included as a published, regulatory-grade precedent for Axiom 6: the optimum of a multivariable process is represented as a bounded region, and a range is a precise claim of equivalence within its bounds rather than an admission of uncertainty. A recommendation of 2-3 minutes of rest between sets is, in this vocabulary, a design space for the rest variable.



