The Problem With Benchmarks
Your beholden standards are holding you back.
In past years I’ve written “fitness standards and benchmarks” to try and describe different tiers of capability across multiple physical domains and energy systems. I haven’t updated those benchmarks in about three years because I realized that the more I tried to categorize the experiences I was trying to covey on a vapid spreadsheet, the more incapable I became.

Old habits and ways of thinking die hard, usually because they served us well at some point — though with varying degrees of consequence. I came across an Instagram / Substack post with strength and conditioning standards reminiscent of what I had previously written.
The post was similar both in categories of movement and tiers of performance. Initially, I “liked” it. Then I hated it.
I don’t know anything about the trainers or organization that published it, so that’s not the point of this article. The point is that I realized the more we’re beholden to any idea or ideal, the more imprisoned by it we become.
The two major complaints I have here are regarding (1) content and (2) process. The first, smaller, problem is the content.
If one were to describe an “elite” standard as:
Bench Press: 1.5x BW
Back Squat: 2x BW
Deadlift: 2.5x BW
Pull Up: 15
5 Mile Run: 35:00
6 Mile Ruck (45 lbs.): 1:15:00
Then we should first note that the time frame of achievement matters greatly. If this is a career resume, it’s not bad, but certainly doesn’t knock my socks off. If we’re talking about a rotating seasonal cycle (e.g. winter strength, summer endurance, and fall capacity) then this is much more impressive, but still doable.
If the “standard that keeps us accountable” — which the post in question asserted — is that one should be able to perform any of those 6 tasks on-demand at any time throughout any given year, then I’m extremely skeptical.
Benchmarks with multiple tiers often come with the implication that scaled (down) tiers are of lower standing (i.e. for novices to aim for) and that the highest tier (RX?) is the “real standard.” What happened to being “cross-fit” enough that physical fitness was not a limiting factor in my skill / sport / or operational tasks?
In a more practical sense, Dan John uses similar “expected” and “game-changer” verbiage for lower tiers 1 and 2, but most importantly denotes the third, highest, tier as “specialized.”
This is much more honest in my opinion, as it implies that exceptional performance in one area will always come at the expense and exclusion of other areas.
The second, much larger, problem is in the process of how we (as an industry and collection of practitioners) define these “standards” and “benchmarks.” In the case above we have what appears to be intended as tests of:
Strength,
Aerobic Capacity, and
Endurance.
First, we have three classic power lifts. Note the important distinction in nomenclature. Strength may be a pre-requisite for expressing power, but it is not tested here — save for the vertical pull (pull ups). Further, all of the “big lifts” involve an axial (balanced) load executed through a bilateral (even plane) movement.
This is where we start to see a breakdown in “functional” and “preparedness” language as well. Synthetic benchmarks and testing are fine, but they are a poor representation of real-world or sporting demands. We all know the difference between “farmer strength” and “gym strength.”
Next, we have the run. The problem here is that running is a skill. In fact it’s not just one Olympic sport, but several of them. This means that if my skill in that sport doesn’t match the capability of my engine (heart and lungs), this is a very poor test of aerobic capacity. On that note, there are far better ways to assess mobility (the structural chassis) than by running.
Speaking of structural chassis, lastly, we have the ruck. If I weigh 180 lbs, and want to be “elite” then I’m expected to keep about a 12:00 / mile pace while carrying 25% of my bodyweight. Perhaps I’m just an idiot, and under-whelmingly not-elite, but that sounds like acute orthopedic trauma to me.
I suspect that the ruck is intended to test endurance, but here’s the catch:
Endurance isn’t just about suffering. It is about learning to love, which is required to “endure” said suffering.
We should also ask what standards are trying to prove, confirm, or question. Otherwise, we succumb to appeals to authority. Is X benchmark “the standard” because Austin said so? Because an Olympian from a bygone era said so? Because someone on a magazine cover or badass on YouTube said so?
Recall the immortal words of Mel Siff:
“Any fool can create a program that is so demanding it would kill an elite athlete, but not any fool can create a tough program that produces progress without unnecessary pain.”
So, what do tests prove? At a minimum it proves that one of two outcomes are true — you either pass or fail. What we learn from those outcomes can be a vastly different story.
If we give a little and pass easily, we likely learn nothing. If you’ve checked all the boxes and got all green lights, you aren’t trying hard enough. When a test (or it’s incumbent measures) become a goal and training in and of themselves, they cease to be a good test (Goodhart’s Law).
On the other hand, to give everything and still fail can be an elucidating experience like no other. Two experiences like this that stand out to me are:
April 2023: 158 assault bike calories in 10:00 (~180 lbs bodyweight).
July 2023: 6-hour AMRAP, 40 bike calories + run 400m (35 rounds total).
The former was an agony I have no desire to replicate, but the lessons have stuck with me ever since. The goal was to hit your bodyweight in calories in 10:00. I failed. Similarly, the latter shattered and redefined the limit of what I thought was possible to “endure”; of no small note, that particular day peaked at 105’F.
Similar to this discussion about benchmarks and standards is the question of “quit or don’t quit.” Immediately, a lot of fitness bros are going to jump on an obvious choice, one that surely reveals their limitations and deficits.
Sure, a lot of people quit early. Those folks are far more concerned with surviving than performing. If you’re struggling to show up, that’s a completely different battle than chasing a standard, let alone setting it.
On the other hand, quite a lot of “high performers” (SOF, CrossFit, combat sports, etc.) have the exact opposite problem. They bury themselves by their own hand thinking it more noble.
The picture where I found the above benchmarks also contained the disclaimer “not intended to be tested on the same day.” Yet, the publisher’s profile prominently displayed the quote “no one cares what you can do fresh.”
That is the take-home irony.
If you’re always half-beat-up you’re always going to get half-ass results, even if they look pretty good relatively compared to the rest of the field. But who and what are we measuring ourselves against? What I believe I am capable of should not change based on what someone else does, but we see this all the time (the 4:00 mile being a prime example).
Meanwhile, you don’t get extra credit for putting an asterisk next to your sliver medal. The pity party of “yeah, but I did it the hard / right way” will never feel as good as gold. Anyone who has been there or even seriously in the hunt knows this in their bones.
“To know your limit is to set one.”
~ Michael Blevins



![The Integrated Fitness Problem [v1.0]](https://substackcdn.com/image/fetch/$s_!jLBX!,w_140,h_140,c_fill,f_auto,q_auto:good,fl_progressive:steep,g_auto/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F222dc620-af54-4649-92fb-720ba44f9c06_922x922.webp)