Follow me on Twitter!


Showing posts with label 2013 Open. Show all posts
Showing posts with label 2013 Open. Show all posts

Sunday, January 4, 2015

How Many Years Do CrossFit Athletes Last in the Open?

What you need to know from this post:
  • About 50% of athletes who compete in the Open on year will continue to the following season.
  • Athletes who have competed for multiple years have a higher likelihood of returning the next season than athletes who are competing in their first season.
  • Based on data from the past four years, approximately 20% of athletes who start competing in the Open will still be competing in their fourth season. I estimate that approximately 10% will still be competing in their seventh season.
  • The higher an athlete ranks in the Open, the higher the probability that they will return the next season.
As I mentioned a few weeks ago, I will be missing the Open for the first time this season.  I had arthroscopic hip surgery to repair a torn labrum.  Without getting into too much detail, the injury was a chronic wear-and-tear type of injury that wasn't a direct result of CrossFit, per se, but rather the fact that I had been so active for the past 10-15 years.  I had a hip impingement that I was born with that put me at risk for this type of injury.  I hope to return to CrossFit eventually, but I won't be ready to compete by late February.

So although my injury isn't directly attributable to CrossFit, it's forced me to face the fact that competing and training at the level I had been is not always easy to sustain.  For me, it was my hip.  For others, it is a shoulder or a knee.  This is true in any sport, and CrossFit is not immune.  This is not a commentary on whether CrossFit as a training methodology is dangerous.  That's a third rail I don't intend to touch.  My point is simply that injuries are inevitable when competing in any serious sport.

With an assist from my wife, who came up with the idea for this post, I decided to try to answer the the question: how long do CrossFit athletes tend to compete?  I'm not talking just about the Rich Fronings of the world, but the everyday athlete who signs up for the Open with no hope of even sniffing Regionals.  We see the Open growing in size each year, but that doesn't necessarily mean athletes are continuing on for multiple years; we could just be replacing the vast majority of athletes one year with an even bigger crop of newbies.

Of course there are multiple reasons an athlete won't compete: injury, lack of interest, disappointment from poor performance in a previous season, work commitments, etc.  With the data I have available, I can't identify which factors are most important, but I can get a pretty good idea of the rate at which athletes are dropping off.

Using the Open data from 2011-2014, I looked at how many athletes continued on each year and what attributes about the athlete (prior Open experience, prior Open finish, age, weight, height) may have an influence on their likelihood to continue.  Because I only have athlete names without any other identifier, this analysis is limited to "truly unique names," which are those that never appeared more than once in any year.  I also focused my analysis on men, because women's names are frequently changed due to marriage or divorce, making it hard to track which athletes truly dropped out and which simply changed name.*

The way I am defining it, an athlete successfully survives from one year to the next if he completes all the Open events in both years.  Failing to complete all the Open events in a given year is counted the same as not competing.  Keep that in mind, since typically about 30% of the Open field from week 1 is gone by the end of the Open.  For purposes of this analysis, those athletes never even competed.  I don't care that HQ still got their money.

Finally, in this study, once you're out, you're out.  If an athlete skips a year, I ignore whether or not they returned the following year.  This makes things much cleaner for the analysis.  In case you are wondering, only about 15% of athletes skip a year and return in a subsequent year (I hope to be one of those 15%!).

OK, so let's get to the results.  First, the simplest way to look at this is to evaluate the athletes that started in 2011 and see how many were left in each subsequent year.  Let's take a look at those results below.


You can see that nearly 60% of athletes survived to Year 2 and nearly 30% were after Year 4.  The issue with this analysis is that it ignores the athletes who started after 2011, which is a huge chunk of the current athlete pool.  The type of athlete who is competing today may be characteristically different than those who started in 2011, and we want to capture that.

The next chart estimates a survival curve for today's athlete population using only the most recent information.  To get the survival rate from Year 1 to Year 2, I looked at athletes who first competed in 2013 and see how many returned in 2014.  To get the survival rate from Year 2 to Year 3, I looked at athletes who competed in 2012 and 2013, then found out how many of those returned in 2014.  To get the survival rate from Year 3 to Year 4, I looked at athletes who competed in 2011-2013, then found out how many of those returned in 2014.

To get the cumulative survival rate for year 3, I multiplied the survival rates for years 1-2 and years 2-3.  To get the cumulative survival rate for year 4, I multiplied the survival rates for years 1-2, years 2-3 and years 3-4.  I then took it one step further, estimating survival rates in years 5-7 based on the rates for the first 4 years.  These are obviously estimates, as we don't have enough data to know the true likelihoods beyond year 4.


Here we see the survival rates are much lower than we previously observed.  About 50% of athletes remain after year 2, about 20% are left after year 4 and I'm estimating only about 10% will be left after year 7.  Clearly the majority of athletes don't make it too many years in this sport, but there is still a decent chunk of the population that sticks with it for years.

However, what's not obvious in the chart is that the chances that an athlete continues in the following year increases with each subsequent year of participation.  Below are the year-to-year survival rates:
  • Year 1-2: 47%
  • Year 2-3: 61%
  • Year 3-4: 72%
This is good news in my opinion.  We see that athletes who stick with it beyond the initial year are not likely to "burn out" the longer they compete.  Once an athlete is sufficiently invested, they are pretty likely to keep at it.**

In total, when you combine first-, second- and third-year competitors, about 52% of the field returns in the following season.  One offshoot of this is that if we consider how fast the Open has been expanding, we know that it must be largely made up of first-time competitors.  In order to maintain the size of the Open field from one year to the next, we need a lot of new athletes each season.  If there were 200,000 total athletes last season, we likely need about 100,000 new athletes to enter the field in 2015 simply to maintain the same size field as before.

The last piece of this analysis was to try to identify other factors (aside from number of years of prior experience) that made certain athletes more or less likely to continue in subsequent years.  For this, I limited the data again to only 2013 and 2014, then further limited the data to athletes who submitted a height and weight (I used the 2013 height/weight for all athletes).

Using a logistic regression model, I looked to see if 2013 percentile rank, age, height or weight had statistically significant impact on the likelihood of returning in 2014.  As it turned out, all but height had a statistically significant impact (p-value less than .01 for percentile rank, age and weight).  However, in my opinion, we can basically ignore weight because the predicted probability of returning did not vary a whole lot (only about 4% higher probability at 180 lbs. vs. 220 lbs.).  

The big key was the 2013 percentile rank.  As you might expect, athletes who finished near the top of the rankings had a much higher likelihood of returning.  Holding all other items at their mean, the predicted probability of returning was 74% for an athlete finishing in the 1st percentile, compared with 55% at the 50th percentile and 35% at the 99th percentile***.  If you aren't convinced by those somewhat opaque predictions from the logistic regression, the chart below shows the observed percentage of athletes who returned, by 2013 percentile rank.




As you can see, the pattern is very evident.  I'm not sure this will come as a surprise to anyone, but it's always nice to see your intuition confirmed in the data.

Age was an interesting factor.  There are two things going on here:
  1. Without controlling for the 2013 percentile rank, it appeared that age had basically no effect.  
  2. Older athletes tend to have worse rankings than younger athletes in general. As we just showed, athletes with lower rankings have lower persistency.
What happens here is that the logistic regression showed that all other things being equal, older athletes actually have a higher likelihood of returning than younger athletes. If we hold the other items at their mean, the predicted probability of returning was 51% for a 20-year-old and 65% for a 50-year-old.

Of course, the big question moving forward is how things will change with the changes made to the Open in 2015.  The addition of a scaled division is likely to siphon off some athletes who had previously competed in the Rx'd division, and it's possible that with only 20 Regional invitations in each region (as compared to 48 in 2013-2014 and 60 in 2011-2012), some additional athletes will drop out.  I also have to believe that the overall participation in the Open will not continue to expand the way it has since 2011, where it has roughly doubled each season.

There is no way to know the answers to these questions right now, but understanding what has happened in the past will certainly help us understand the impact of these changes in the future.

[Thanks a lot to Andrew Havko, Michael Girdley and Jeff King for pulling this data for me and/or making it publicly available]

*It appears that overall, the survival rate for women is very similar to that of men.  Initially, it appears about 4% lower, but using some very rough data for the marriage and divorce rates, that discrepancy could very easily be attributable entirely to name changes.

**For my estimates beyond year 4, I assumed this would continue to flatten out, so I used 77% for year 4-5, 79% for year 5-6 and 81% for year 6-7.

***These predicted probabilities are all probably about 4% too high.  That's because the subset of athletes used for the logistic regression was limited to those that submitted a legitimate height and weight.  In general, these athletes have slightly higher finishes and are slightly more likely to return than the average athlete.

Saturday, December 13, 2014

Understanding Year-to-Year Improvement in the Open

In this and most future posts, I'll start by giving you a very brief summary of the key takeaways from the post.

What you need to know from this post:

  • On average, athletes that compete in the Open in multiple years improve their rank percentile in the second year, but their absolute ranking declines.
  • The more years an athlete competes in the Open, the less they improve their percentile ranking in subsequent years, although the average improvement is still positive. 
  • There is pretty strong evidence that if the Open has a higher load-based emphasis on lifting (LBEL), this will favor taller and heavier athletes. If the Open has a lower LBEL, this will favor smaller and lighter athletes.
  • It is unclear how age is related to an athlete's percentile ranking improvement from year-to-year.

Today I want to get into a topic that has interested me for some time.  Many readers of this site have been competing in the Open for several years and likely have used their performance each year to judge how much their fitness has improved from the past year.  There are two underlying assumptions that we use here:
  1. The Open is a pretty good test of overall fitness;
  2. Each year of the Open is a relatively similar test of fitness, compared to other years.
I'm going to leave the first assumption unchallenged today, although there could certainly be debate about that.  But let's assume that the Open is indeed a good test of fitness.  What I will try to do today, using data from 2011-2014, is try to test the second assumption and get a feel for how much impact, if any, variations in the programming might have from year-to-year.  Along the way, I'll also look for other interesting observations about how athletes are improving across multiple years in the Open.

Before we get into the results, here's a quick background on the data I'm using and my basic methodology:
  • For both men and women, I started with the Open results for athletes under the age of 55 (meaning no scaling in the Open).
  • I removed any athletes whose first/last name combination was not unique.  This is due to the fact that my data does not have any other identifiers for each athlete.  Since there are something like 9 or 10 Ben Smith's, I just threw them all out.
  • I removed all athletes that did not complete all five events.
Next, I split up the analysis into six cohorts: 2013-to-2014 male, 2013-to-2014 female, 2012-to-2013 male, 2012-to-2013 female, 2011-to-2012 male and 2011-2012 female (an athlete could be multiple cohorts).  For each section, I identified athletes that competed in both years.  For all athletes that submitted it, I also mapped on age, height and weight information from 2013 (except for the 2011-2012 cohorts, in which case I used the 2012 information).  I had to make the simplifying assumption that an athlete's weight did not change from year-to-year, which is probably not true for some athletes.

OK, with the background out of the way, let's move onto the findings.  The first thing I wanted to know was just how much improvement athletes were making from year-to-year.  The initial results might surprise you:


On average, athletes who continued from one year to the next actually finished lower in the second year than in the first.  In fact, between 60-80% of athletes had a lower rank in the subsequent year across all six cohorts.  

So what gives?  Well, the key here is that the field has been expanding, nearly doubling in size each year.  The easy way to account for that is to look at the change in an athlete's percentile rank from year-to-year.  If an athlete is 5,000th out of 50,000 in 2013 and 8,000th out of 100,000 in 2014, then the percentile rank actually shows an increase of 1% (5% to 4%), despite the 3,000-spot drop in absolute rank.


Now we see that in general, athletes are improving year-to-year, although it became a bit more difficult each year.  Approximately 89% of athletes improved their percentile rank from 2011-2012, compared to 80% from 2012-to-2013 and 71% from 2013-to-2014.  

Each year, we have more and more athletes who have been competing for several years.  Most of us who have been CrossFitting for a long time know that making incremental improvements becomes harder and harder (thought not impossible) as the years go on.  For evidence of this, I looked at the average percentile improvement from 2013-to-2014 of athletes who also competed in 2012 vs. those who did not.*


These numbers make it fairly clear that the amount of past Open experience is a factor in how much improvement athletes make from year-to-year in the Open.  But how about other variables, such as height, weight or age?  This is where things get a little tricky.

First, let's look at age.  The three charts below show the average improvement by age for males (orange) and females (blue) in each time period.




From these charts, we see that there is not a simple answer here.  In two cases (male 2012-2013 and female 2013-2014), improvements were generally higher at older ages, but in the other four cases, the reverse was true.  Across the four cohorts, the correlation between age and percentile rank improvement ranged from -12% to +14%.  Unfortunately, the results are not consistent by gender or by year, in which case we might be able to make some sort of generalization about what these results mean.  For now, I will simply conclude that there is no clear relationship between age and year-to-year improvement in the Open.

How about height and weight**?  Well, again, the results were mixed, but in this case, the mix of results might actually be able to provide some insight.  Let's focus on weight for now.  Four of the six cohorts did show a clear linear relationship between weight and percentile rank improvement, but the other two did not.  Here are charts showing the relationships for those cohorts.


From 2012-to-2013 for females (third chart), it appears that heavier athletes tended to show more improvement than lighter female athletes.  In the other three cohorts shown, heavier athletes tended to show less improvement than lighter athletes.  

Another way to look at this is by examining the correlation between and percentile rank improvement  in each cohort.  A positive correlation means that higher weights tend to have higher percentile rank improvements; a negative correlation means that higher weights tend to have lower percentile rank improvements.  The correlations for the four cohorts shown above tell the same story: -11% correlation for 2011-2012 males, -15% correlation for 2011-2012 females, +8% correlation for 2012-2013 females and -6% correlation for 2013-2014 males.

What could be the reason for these results? 

For each of the Opens, I've evaluated the programming using a few metrics.  One that I reference quite frequently is the load-based emphasis on lifting (LBEL).  This metric attempts to quantify how "heavy" a CrossFit competition is based on the portion of the competition that was made up of lifts (as opposed to bodyweight movements), as well as how "heavy" those movements were.  After seeing the charts above, I went back and looked at the LBEL in each of the years for both males and females.  What I found was that in situations where there was a strong negative correlation between weight and improvement, the LBEL decreased significantly, and in situations where there was a strong positive correlation between weight and improvement, the LBEL increased significantly.

The chart below shows the following for each cohort:
  • Percentage change in LBEL;
  • Correlation between weight and improvement in percentile rank;
  • Correlation between height and improvement in percentile rank; and
  • Correlation between age and improvement in percentile rank.

While we still can't seem to tell much about the relationship between age and improvement, this chart above does seem to clearly indicate that a higher LBEL benefits larger athletes and a lower LBEL benefits smaller athletes.  Note that the color scales for the percent change in LBEL mirror the weight correlation almost exactly.

While this my seem intuitive, I think it is a very important result to lend credibility to the LBEL metric.  It is still my belief that LBEL is not a particularly useful metric when looking at individual events, but when used to evaluate a multi-event competition such as the Open or Regionals, I think it is very useful to help us understand the type of athletes that might benefit from the programming.  Keep in mind, of course, that these correlations are relatively small, so there are plenty of other factors that determine how much an athlete will improve from year-to-year.

Obviously, none of this is meant to de-emphasize the importance of training in determining how much an athlete will improve from year-to-year.  Rather, this can help us understand all the factors that might be impacting an athlete's improvement, which can in turn help evaluate how successful all that training really was.

[Thanks a lot to Andrew Havko, Michael Girdley and Jeff King for pulling this data for me and/or making it publicly available]

*I did also take a look briefly at the 2013-2014 improvement for athletes who competed back in 2011.  As expected, they were slightly lower than those who just competed in 2012.  However, they did still show a positive improvement in their percentile rank, on average.

**Many athletes did not submit their height or weight (or typed in something ridiculous, like 1,000 pounds).  Any time I looked at correlations between weight/height and percentile rank improvement, these are based only on the subset of athletes that reported a reasonable height and weight.  This ranged from about 50%-80% of the field (women generally reported less often).


Thursday, February 13, 2014

Height, Weight and Programming: Part 1

For a long time, I've stayed away from investigating which height, weight, age, BMI or other characteristics yield the best results in CrossFit. Part of the reason is that there have been several other blogs that have already put together some nice plots and analyses about these topics. The other reason is that I wasn't sure how much it helps the average CrossFitter to know these things. For instance, if you find that your height is not "ideal" for CrossFit, how does this help you? You can't change your height, and obviously I wouldn't encourage you to stop competing because your current height might put you at a disadvantage. Even with weight, I think it's generally not worth trying too hard to change your weight unless you are drastically over- or underweight. Just keep training, eat well, and I believe your body will work its way into the weight that works for you.

However, I started thinking that there is value in looking into these topics. CrossFit does not have weight or height classes, and so I think it's important for the sport as a whole that the programming be as unbiased as possible. It seems that this has been a constant challenge since the inception of the CrossFit Games, and likely something Dave Castro and the folks at HQ are considering as they put together the workouts for each competition. What I'm hoping to do here is look at the data and see how well that goal is being accomplished.

One thing to consider, however, is that for the sport to be as fair as possible to all sizes of individuals, it's not imperative that EVERY workout be unbiased. In fact, one of the thigns that makes the sport so fascinating to me is seeing the bigger athletes like Aja Barto or Chad Mackay tackle the supposed "little guy" workouts and seeing the Chris Speallers of the world being forced to push a 400-lb. sled. One of the ways I've attempted to quantify the bias of individual workouts, as well as entire competitions, is using the concepts of average relative loads and load-based emphasis on lifting (LBEL). See my post "What to Expect From the 2013 Open and Beyond" for a full explanation, but essentially the average relative load tells us how "heavy" the weights were across a workout or a competition, and the LBEL tells us how much emphasis there was on lifting, with heavier weights getting more value.

My theory has been that the higher the LBEL, the more the competition should favor a bigger athlete. I believe this true for an individual workout, but as we will see, these metrics aren't always that precise when we look at just one workout. One reason is that the metrics assume each movement in a workout is worth equal value. Broadly, this is true if we look at several workouts together, but certain movements aren't programmed so that each movement is truly valued equally. For instance, on 12.4/13.3, it's fair to assume the 150 wall balls were worth far more than the 90 double-unders, since the double-unders can be completed in 1-2 minutes for a competitive athlete, while the wall balls take 5-6 minutes for most top athletes. Additionally, the muscle-ups don't even come into play for roughly half the field.

Today, I'm going to look solely at the 2013 Open results and see what we can learn (data again provided by Michael Girdley at http://girdley.com). The first thing I've done is to look at the relationship between weight and performance on each 2013 Open event. However, in doing this, I tried to normalize for the BMI of the athletes at each weight level. The way I did this is by calculating the average finish for athletes in each bodyweight/BMI combination, then calculating a "weighted" average for each bodyweight class, with the "weights" for the weighted average based on the mix of BMI's across the entire field.

The reason is that the heavier athletes may include more athletes who are overweight and perhaps not in great shape. However, what I'm interested in is comparing athletes who have roughly the same general fitness level but differ in terms of weight (think Jason Khalipa vs. Chris Spealler). BMI is far from a great indicator of fitness for an individual person, but it's relatively unbiased, and it will help us level the playing field a bit. Ideally, something like body fat percentage of V02 max would be a good way to normalize, but we don't have access to that information yet.

Below are two graphs (one male, one female) showing the average ranking for athletes by weight on each of the 2013 Open events. Note that these are based only on athletes under 40 years old who finished all five events and had a height and weight within a reasonable range. I limited the field to this group of athletes and then re-ranked them on each event before performing this analysis.



The thing that stands out to me here is that all of the first four events followed a similar pattern for both men and women, while the fifth event followed a distinctly different pattern. The first four events had an ideal weight somewhere between 185-195 for men and 150-160 for women. The fifth event, however, heavily favored the smaller athletes. For the women in particular, the graph never "bottomed out," meaning that essentially the lighter the athlete, the better the finish (due to sample size, I really couldn't draw any reasonable conclusions about weights outside the range shown).

Does that make 13.5 a bad event? Not necessarily. But it does mean that the chest-to-bar pull-ups appear to have been more important than the thrusters, considering the event generally favored smaller athletes. If that was the intention, then there's no issue.

Now let's look at the same analysis for height.



Here we see a similar pattern. All the first four events had a similar "sweet spot," with the fifth event favoring a much smaller athlete.

However, what may surprise you (it surprised me) is that the ideal height for men and women is surprisingly close. For men, the ideal height for those first four events ranged from about 5-11 to 6-1, while for women it ranged from about 5-9 to 5-11. Looking only at the men's results, one might assume that the ideal CrossFit athlete is one who is average height; looking only at the women's results, one might assume that the ideal CrossFit athlete is one who is taller than average. However, it may actually be that CrossFit tends to favor athletes who are near 5-10, regardless of gender.

The chart below summarizes the findings by looking at the ideal height and weight for each event, and for the competition as a whole. The "ideal" here is not simply the one with the lowest rank, but rather a weighted average of the 3 heights/weights with the lowest ranks, with more weight given to the height/weight with the absolute best rank. This helps smooth out our results a bit.


What we do not see here is much correlation between LBEL and the ideal height or weight. The event that favored the smaller athletes was 13.5, which had an LBEL right in the middle of the pack, while 13.3, the lightest event, was ideal for larger athletes. But as mentioned above, for 13.3, the LBEL may be deceptively low due to the design of the workout.

So can we assume that LBEL tells us nothing about which athletes each events favor? I think more study is needed. For one thing, looking only at the ideal heights and weights do not account for how much the smaller or larger athletes are penalized. Also, the 2013 Open gave us pretty homogenous events: no one event was particularly heavy or particularly light. In the 2012 Open, on the other hand, we had an unweighted event (12.1) and a lifting-only event (12.2). Perhaps a future post can run this same type of analysis on the 2012 Open, although there has already been some work done on that http://xfit2011.blogspot.com/ and https://sites.google.com/site/cfopen2012analysis/home. Also, for me to do my normalization by BMI, the smaller sample size in 2012 could pose a problem, particularly for the women.

But as I mentioned earlier, I expect that LBEL is more informative when comparing entire competitions (for instance, regionals vs. Open) than when comparing individual events. One thing I plan to investigate in a future post is athletes of varying heights and weights fared at regionals, accounting for how well they performed in the Open. Given that the LBEL was much higher at the 2013 Regionals than at the 2013 Open, we would expect that for two athletes who performed equally well in the Open, the larger athlete would have an advantage at regionals. We shall see whether the data confirms this.

So what are the big takeaways today?




  1. Overall, the 2013 CrossFit Open favored men around 5-11, 190 lbs and women around 5-10, 155 lbs. I think it is likely that these are roughly the ideal heights and weights for CrossFit in general. For women, this may come as a surprise that the ideal athlete is so tall.
  2. The advantages toward any particular height or weight are minimal in total. Athletes at the ideal weight finished only about 4,000 spots ahead of athletes at the least ideal weight (out of 40,000). For women, the largest gap was about 2,000 spots (out of about 20,000). The same was true for height.
  3. The first four events in the 2013 Open favored athletes around the overall ideal, while the fifth event favored much smaller athletes.
  4. For individual events in the Open, a higher LBEL doesn't always mean the event favors larger athletes. However, it is still possible (and likely in my opinion) that that a higher LBEL across an entire competition means the competition favors larger athletes. More study is needed, and I'm hoping to dive into that after this year's Open.
Hopefully, today's post provides a good starting point for this discussion. But to be sure, there is still more work to be done.


In other news, the Open starts in 15 days, so get your SWAG's ready for event 14.1! From here until the end of the Open, that's going to be the focus of my posts on here. My goal is to get 2 posts regarding each event, but cut me some slack - I've got a 3-month-old baby and training of my own, but I'll be doing the best I can.

Until then, good luck with your training!

Friday, December 20, 2013

Are CrossFitters Specializing?

For the past seven years, the CrossFit Games have sought to find the fittest all-around athletes in the world. One common theme from CrossFit HQ is that the Games seek to "punish the specialist" and "reward the generalist." However, one can't help but notice the increased attention that certain aspects of fitness seem to garner in the CrossFit community compared to others. In CrossFit media, the emphasis on lifting, and Olympic lifting in particular, seems to be disproportionate to many other areas of fitness. When was the last time you saw a video or even a note on the Games site mentioning a Games athlete hitting a new PR in his 5K run? Yet it seems like it hasn't been 24 hours since we've seen a video posted of another athlete hitting a big snatch or clean and jerk.

OPT noted this phenomenon on his blog about a year ago:

"Media recently for the sport has put an emphasis on strength development in spite of promoting true “balance” in fitness and the general components of fitness.  A sport where now the elite can qualify for the American open weightlifting championships but cannot qualify for a state-level high school cross country meet."

Although I won't seek to prove it in today's analysis, I think there is little doubt that the big lifts tend to get a lot more attention than general metabolic conditioning in CrossFit media. However, the question I will attempt to answer is whether the sport itself has gotten out of balance.

From the programming perspective, I showed in my recent post "History Lesson: An Objective, Analytical Look at the Evolution of the CrossFit Games" that while the metcons have become heavier and heavier over time, the overall balance of lifting and conditioning has not changed drastically in the past 7 years. In addition, there is roughly the same amount of emphasis on Olympic lifting now as there has been throughout Games history. In fact, there has actually been a shift away from the powerlifting-style movements like the deadlift and back squat. Running is and has always been the most common movement at the CrossFit Games.

However, there is a legitimate question about the intense focus on Olympic lifting and the lack of focus on running at the Regional level. And there is also no doubt that the CrossFit athletes have been getting more and more proficient at the Olympic lifts, as evidenced by the rising numbers in the 1-rep max events at the Games each year.

So let's ignore the programming for now and focus on the actual strengths and weaknesses of the athletes in our community. Before I do this, I want to re-visit the comment from OPT above. I love OPT and have probably watched every CrossFit.com video of his over the past few years, but I think the commentary that "the elite can qualify for the American open weightlifting championships but cannot qualify for a state-level high school cross country meet" is a bit misleading on a couple of levels:
  1. The American Open is not that competitive on a global scale. To qualify in the 85 kg weight class, you need a 266 KG total (http://0205632.netsolhost.com/2013NationalEventsQualifyingTotals.pdf). Yet at the 2012 Olympics, the 16th place finisher had a total of 315, or 18% higher. On the other hand, to qualify for the highest-level state cross country meet in Ohio (a competitive state where I used to cover sports), you need a time around 17:00. Considering these meets are run on rugged terrain rather than on a track, this time isn't that far behind the Olympic 5,000 meter times (13:52 was 15th place in the 2012 Olympic final). So I think it could be argued that qualifying for those two events are actually relatively comparable as far as difficulty.
  2. There are weight classes in Olympic weightlifting, yet there are not in CrossFit or cross-country. The top CrossFitters are not even in the same stratosphere as the lifters in the 105KG+ weight class. On the flip side, among runners above 200 lbs., I have to believe someone like Garrett Fisher would be considered elite.
But let's look at the numbers throughout our community. To do this analysis, I used the 2013 Open data, which was generously pulled and cleaned for me by Michael Girdley (girdley.com). This dataset has all the numerical information provided by athletes who competed in the Open (it does not include answers to the questions about diet, how long you've done CrossFit, etc.). Based on this self-reported data, I believe we can understand how CrossFit athletes from top-to-bottom compare to the world's best in a variety of lifts, running events and metcons.

To perform this analysis, I first limited the data to athletes under 40 who completed all five events (approximately 39,000 men and 23,000 women). Then I re-ranked all the athletes based on their rank across all 5 Open events and grouped them into 20 buckets based on this rank. Within each bucket, I took the average for each of the self-reported scores (Fran, Helen, Grace, Filthy 50, FGB, 400 meter run, 5K run, clean and jerk, snatch, deadlift, back squat, max pull-ups). For the timed events, I converted these to a pace (rounds per second, for instance, or meters per second). Then, I pulled in world records for each and compared the CrossFit community against those world records.

The charts below shows how the community compares to the world records*. For the lifting events, these are the world records without regard to weight class, since CrossFit does not have weight classes. To reduce clutter, I have grouped all metcons together, all runs together and all lifts together (pull-ups stayed in their own category).





On both charts, we can see that the community is generally closest to the world record when it comes to running events and (not surprisingly) metcons. When compared to the world record in the lifts, it becomes obvious how far behind the CrossFit world still is. Proud of that 200-lb. snatch? Congratulations, you are slightly below half the world record. (Note: I am proud of my 200-lb. snatch, and it took me 5 years to finally get there).

But let's also look at the Games athletes in particular and see how they stack up. Here is a table showing each event and how the average Games athlete stacks up compared to the world record.


Not shockingly, the Games athletes are near the top to the world record in the metcons (since generally the world record comes from this field), but again, note that they are much closer to the world records in the 5K run and the 400 meter run than they are in the Olympic lifts. And look at that back squat - not even close! (and before you ask, this is comparing against the raw world record in the back squat, which appears to be 450 KG as best I could tell)

It is interesting to note that the elite CrossFit men are closer to the Olympic lifting world records than the elite women, yet they are further from the 5K run record. This could have something to do with the background that many of the athletes had prior to CrossFit, but that's purely a guess at this point.

Another way to look at this is to understand where the Games athletes excel furthest beyond the rest of the community, and in particular, where they excel furthest beyond the rest of the Regional field. For both men and women, here is a look at how the top 5% of Open finishers (roughly the Regional field) compared to the Games athletes.


Here I think we start to see something interesting. While the Games athletes do not appear to be any further from the world record in the 5K run than they are in the Olympic lifts, they aren't that much better than the rest of the regional field when it comes to the running events. They also aren't that much better when it comes to the powerlifting movements. I think you could attribute at least partially to programming at the Regionals: we simply aren't testing much for running or powerlifting, so the athletes making the Games aren't necessarily that much better than the rest of the field in those areas.

On the flip side, look where the Games athletes do exceed their peers by a greater amount: the Olympic lift, the short metcons and pull-ups. It seems that explosive power and conditioning (over a relatively short time frame) are what tend to separate the Games athletes from the rest of the Regional field.

One last way to look at this is to see the gap between the Regional athletes and the median Open athlete**, which is defined as the athletes finishing in 45th-55th percentile in the Open among people under 40 who completed all 5 events. These median Open athletes are still generally fit individuals, they just aren't quite at the Regional level.


This table looks a lot like the prior one, meaning that what separates the Regional athletes from the average Open athletes is a lot like what separates the Games athletes from the Regional athletes.

Based on the analysis here, I believe that CrossFit athletes in general aren't bad runners or particularly tremendous lifters. However, the elite CrossFit athletes are significantly better lifters than the rest of the community, and yet they are not drastically better runners than the rest of the community. From this perspective, we do see a little bit of the bias that OPT was writing about. But overall, I don't think the specialization issue is as much of a concern as some might think.

Update 12/20: [I'd like to note that I don't believe that achieving 70% of the world record in the snatch is exactly as challenging as achieving 70% of the world record pace in a 400 meter run. However, the fact that CrossFit Games athletes are so much closer to the world record in the running events than they are in the lifts indicates to me that these athletes should not be considered specialists in the Olympic lifts who simply neglect running. One way to quantify this, which I'm hoping to look into more, is to put things in terms of standard deviations. I have looked at this for the snatch, clean & jerk, 5K run and 400 meter run for men, however. Using the standard deviation based on the same sample of Open athletes under 40, the Games athletes are approximately 6.0 standard deviations below the world record in the lifts but only 4.3 standard deviations away in the 5K run and 2.0 standard deviations away in the 400 meter run. This isn't a perfect method either, but again, it supports the idea that Games athletes aren't totally specializing in the Olympic lifts while neglecting their running.

However, I do see the same pattern as in the main body of my post when comparing Games athletes to the rest of the CrossFit field. Games athletes are only about 1.1 standard deviations better than the median in the 400 meter run, 0.7 standard deviations better in the 5K run, but approximately 2.5 standard deviations better in the Olympic lifts. So it seems that the same conclusions generally hold when doing the analysis this way.]


Update 12/21: [As a follow-up to the previous update, I looked at where the average Games athlete would fall in the spectrum of all Open athletes in each of the self-reported metrics. This was more difficult than it might seem because of the tremendous selection bias in the data (only about 20% of the men's field reported a 400m time, for instance, but about 50% reported a deadlift max). I tried to account for this by creating a "weighted" distribution, where each 5% bucket was only worth the same number of total athletes, regardless of how many missing values they had. After doing this, I found that the average male Games athlete is in the top 2% in the clean and jerk and snatch, and they were at least the top 6% in all other lifts or metcons. However, for the 400 meter sprint, they were only in the top 15%, and in the 5K run, they were only in the top 25%.

Note that the selection bias still can't be totally accounted for. It's probably fair to assume that the non-responders in general had worse scores than those that did respond, so maybe the Games athletes are actually even better than they appear here. However, it is definitely striking that the Games athletes are not that far beyond their peers in the runs, particularly the 5K. Still, it doesn't necessarily say they are bad runners, as you could argue that CrossFitters in general are good runners and therefore the Games athletes are still pretty good. It is clear, however, that Games athletes are outdistancing their peers substantially in the Olympic lifts, even though they are generally still well short of elite status. 

I think a lot of this has to do with the background of many CrossFitters. Many, many people ran to stay in shape prior to finding CrossFit, but relatively few Olympic lifted. To really decide if you feel the sport of CrossFit has gotten too specialized, I think all of the preceding analysis has to be taken in together, including how far CrossFitters are from the world records as well as how the Games athletes compare to the rest of the field. I'm not sure there is really a clear-cut answer.]

Update 12/28: [Quick one here. I did the same analysis for the women that I did for the men on 12/21, and I found that the women's Games were slightly more dominant across the board. The average Games athlete would be in the top 3% for all lifts and metcons, the top 9% for the 5K run and the top 15% for the 400 meter sprint. Interesting that they were comparatively better than the men in the 5K, although actually about the same in the 400 meter sprint. I'm not sure I really have a good hypothesis for this at the moment.

Also, worth noting is that I looked into the response rates for each metric, and found that for women, the rate was between 7% and 12% for all runs and metcons, except Fran, which was 17%. The response rate was between 32% and 38% for all the lifts and 16% for max pull-ups.

For men, the rates were between 13-23% for all metcons, except Fran, which was 33%. The response rate was between 44% and 51% for all the lifts and 30% for max pull-ups.

This does indicate that there is a selection bias issue that has to be considered, but it's not as if it ONLY applies to the runs. Basically all the metcons and the runs had very low response rates, but the lifts had much higher response rates.]

*Here are the world records I used in this analysis, based on a combination of web research and self-reported PRs from the database: 
Fran - 2:00 (men), 2:07 (women)
Helen - 6:13, 7:20
Grace - 1:14, 1:17
Filthy 50 - 14:05, 16:13
FGB - 520, 460
400 meters - :44, :50
5,000 meters - 12:37, 14:11
Clean & Jerk - 263 KG, 190 KG
Snatch - 214 KG, 151 KG
Deadlift - 461 KG, 264 KG
Back squat (raw) - 450 KG, 280 KG
Max pull-ups - 106, 80

**There is a significant amount of selection bias in these self-reported numbers, which is why I used the bucketing approach to account for it. In general, the people reporting their numbers for each lift/run/metcon are better at those lifts/runs/metcons than those who leave them blank. Also, for many of the metcons, less experienced athletes may not even have a PR. As an example of this bias, if you take a straight average of the clean and jerk across all women under 40 finishing all 5 events, it's about 134 pounds. But if you group the field by the 5% buckets as I have, take the average in each bucket, then average across all buckets, you get an average of 126 pounds, which I believe is more representative of the "true" average.

Wednesday, August 14, 2013

A Closer Look at the 2013 Games Season Programming

I struggled for the last few days on how to present this analysis. Last year, I wrote two lengthy posts assessing the programming for the 2012 Games season. I titled the posts "Were the Games Well-Programmed." While I thought those posts turned out well, I hesitated to simply follow the same template as last year, for a couple reasons:
  • Plenty of people have an opinion on the Games programming, many of whom are much more known in the CrossFit community than me (for instance, I've already read analysis from Rudy Nielsen and Ben Bergeron). Do we need more opinions out there?
  • Assigning grades or giving a thumbs-up/thumbs-down to the Games programming gives off the impression that I have it all figured out. I think HQ has made it clear that they work very hard not to be influenced by the outside world in their decision-making. Am I really going to accomplish anything by telling them they were wrong?
However, balancing those concerns was my feeling that I do have something unique to provide to the discussions. And, most importantly, I think the discussion is important. While I respect HQ's stance to do things their own way, I'd like to think that they are always looking for ways to improve the Games. Although I don't work for HQ, I don't feel as though I'm an outsider. Those of us in the community, and especially those who've been following and competing in the sport for years, are all working toward the same goal: to keep this sport progressing in the right direction. I know that HQ is at least marginally aware of this site, considering Tony Budding took the time to comment on my scoring system post last year. Here's to hoping they're still keeping up with me (and I promise I'll leave the scoring system out of the debate for now, Tony).

With that in mind, this post will be broken down in much the same way as last year's discussion. There are five goals I think that should be driving the programming of the Games, in order of importance:
  1. Ensure that the fittest athletes win the overall championship
  2. Make the competition as fair as possible for all athletes involved
  3. Test events across broad time and modal domains (i.e., stay in keeping with CrossFit's general definition of fitness)
  4. Balance the time and modal domains so that no elements are weighted too heavily
  5. Make the event enjoyable for the spectators
What I'd like to do is assess how well those five goals were accomplished this season. Unlike last year, however, I'm making a couple changes.
  • This year, I'm going to take the entire season into account in this post (last year I separated the Games programming specifically from the Games season as a whole). I've already covered the 2013 Open and Regional programming to some degree in previous posts, so I'll be incorporating some of that here. I think it's better to try to view the Games in the context of the whole season.
  • I won't be giving grades for each goal this year. Instead, I'll be pointing out suggestions for improvement, because simply identifying the problems only gets us halfway there. Additionally, I'll point out things that I felt worked out particularly well. Every year, HQ does a few things that bug me, but they also do a handful of things that make me say, "Hey, that was a great idea. I wouldn't have thought of that." I think it's worth acknowledging both sides.
So with that as our background, let's get started.

1. Ensure that the fittest athletes win the overall championship

I think it's hard to argue this wasn't accomplished this year. Rich Froning was challenged, but he still came out of the weekend looking pretty unbeatable. Sam Briggs, although she did show a few weaknesses, appeared to be the most-well rounded athlete across the board by the end of the weekend, while many of the women who were expected to be her top competition had major hiccups. Both Froning and Briggs won the Open and finished near or at the top in the cross-Regional comparison.

Additionally, as I pointed out in my last post, the athletes that we expected to be at the top generally finished that way. That doesn't absolutely mean that the Games are a perfect test, but it does provide some validation when the top athletes keep showing up near the top across a variety of tests in successive years.

How We Can Do Better: I don't really have anything here. The right athletes won, so mission accomplished.
Credit Where Credit is Due: The fact that almost all the athletes competed in every event really helped keep things interesting until the end. In the past, we've seen athletes build an early lead and hang on simply because the field gets so small that there aren't enough points to be lost in the late events. Allowing 30 athletes to finish the weekend allowed some big swings at the end, including Lindsey Valenzuela's move from 5th to 2nd in the final two events.


2. Make the competition as fair as possible for all athletes involved

Because I promised Tony Budding I wouldn't bring up the scoring system in general, I won't touch on that here. Let's just say I think the scoring system is fair enough. However, the way the scoring system was applied in Cinco 1 and 2 didn't make a whole lot of sense. Any athlete who didn't finish the handstand walk (Cinco 1) or the lunges (Cinco 2) was locked in a tie, despite the fact that the lunges took 2-4 minutes and the separation was very clear between many athletes who were tied. Because of the massive logjam (21 male athletes tied for 7th, 13 female athletes tied for 4th), the few athletes who did finish didn't get that big of a point spread on many other athletes who were on pace to be several minutes behind.

The other issue here is judging, which does tie in with programming to some extent. I think the judging continues to improve each year. Anyone who's been to a local competition has seen the judges who just don't have the stones to call a no-rep. That simply doesn't happen at the Games. You cannot get away with cheating reps, and that's definitely a good thing for the sport.

I won't dwell on it here, but everyone knows the judging in the Open is still a concern (see 13.2 Josh Golden/Danielle Sidell fiasco this year). Hopefully some careful programming will alleviate that next year.

How We Can Do BetterImprove tiebreakers for movements such as walking lunges, handstand walks, running, or anything where a distance is involved instead of a number of reps. Also, I'd prefer to have Games athletes not perform chin-to-bar pull-ups. They are really tricky to judge and aren't as impressive to spectators. In fact, the whole "2007" event just didn't really work for me; it seemed like basically a pull-up contest for the athletes at this level.
Credit Where Credit is Due: Chip timing helped identify the winners really nicely in some of the shorter events. Also, judging keeps improving each year.


3. Test events across broad time and modal domains (i.e., stay in keeping with CrossFit's general definition of fitness)

Right off the bat, let's look at a list of all the movements used this season, along with the movement subcategory I've placed each one into. I realize the subcategories are subjective, and an argument could be made to shift a few movements around or create a new subcategory. In general, I think this is a decent organizational scheme (and I've used it in the past), but I'm open to suggestions.


It's pretty clear that the CrossFit Games season is testing a very wide variety of movements, and the majority of those were used in the Games. Even some that were left out of the Games, like ordinary burpees* and unweighted pistols, were used in other forms (wall burpees*, weighted pistol). No major movements that we've seen in the past were left out of this entire season, with the exception of back squats. I've seen some suggestions online about testing a max back or front squat in the future, as opposed to the Olympic lifts that we have been seeing a lot.

Another key goal is to hit a wide variety of time domains and weight loads. Below are charts showing the distribution of the times and the relative weight loads (for men) this season. The explanation behind the relative weight loads can be found in my post "What to Expect From the 2013 Open and Beyond." Two notes: 1) some of the Regional and Games movements had to be estimated because I don't have any data on them (such as weighted overhead lunge and pig flips); 2) the time domains for workouts that weren't AMRAP were rough estimates of the average finishing times.


Although most of the times were under 30 minutes, we did see a couple beyond that, including one over an hour (the half-marathon row). As for the weight loads, we saw quite a range as well. The two heaviest loads were from the max effort lifts (3RM OHS and the C&J Ladder), but there were also some very heavy lifts used in metcons, mainly in the Games (405-lb. deadlifts for crying out loud). Still, lighter loads were tested frequently in early stages of competition (Jackie, 13.2, 13.3).

How We Can Do BetterI like the idea of testing a max effort on something other than an Olympic lift.
Credit Where Credit is Due: Nice distribution of time domains, and no areas of fitness were left neglected entirely. CrossFit haters can't point to many things and say 'But I bet those guys can't do X.' Yeah, they probably can.


4. Balance the time and modal domains so that no elements are weighted too heavily

Based on the subcategories of movements I've defined above, let's look at a the breakdown of movements in each segment of the 2013 Games Season. These percentages are based on the weight each movement was given in each workout, not simply the number of times the movement occurred (for example, the chest-to-bar pull-ups were worth 0.50 events in Open 13.5, but they were worth only 0.25 events in Regional Event 4).


One thing that surprised me was how little focus there was at the Games on basic gymnastics (pull-ups, push-ups, toes-to-bar, etc.). However, there was quite a bit of bodyweight emphasis (high-skill gymnastics like muscle-ups and HSPU), as well as some twists on other bodyweight movements (wall burpee, weighted GHD sit-up). Overall, bodyweight movements (including rowing) were worth 60% of the points and lifts were worth 40%.

Another surprising thing was how much emphasis there was on the pure conditioning movements like rowing and running. Now, one of the "running" events was the zig-zag sprint, which wasn't actually about conditioning but rather explosive speed and agility. Still, the burden run and the two rowing events really put a big focus on metabolic engine and stamina. I have no problem with this, but what I would like to see is these areas tested more early on. Running in the Open is almost impossible, but at the Regional level, it would make sense to test some sort of middle- or long-distance runs so that athletes who struggle there would have those weaknesses exposed.

As far as loading is concerned, what seems to be happening at the Games in recent years is that things are either super-heavy or super-light. Only two of 12 events tested what I would consider medium loads (somewhere around a 1.0 relative weight for men, like 135-lb. cleans or 95-lb. thrusters), and none tested light loads. Also, as noted above, the bodyweight movements that were required were generally extremely challenging. I personally wouldn't mind seeing some more "classic" CrossFit workouts involved, like we saw with "The Girls" at the end of last year's Games.

Whereas last year's Games seemed to be lacking in the moderately long time frame (12:00-25:00), I think they did a better job of spreading things out this season. In the Games, we had 1 event over 40:00, 3 between 12:00 and 40:00, 4 between 1:00 and 15:00 and 2 that were essentially 0 time.

One other way to see if we're not weighting one area too much is to look at the rank correlations between the events. If the rankings for two separate events are highly correlated, it indicates that we may be over-emphasizing one particular area. For this analysis, I focused only on the Games, because it's not really such a bad thing if we test the same thing in two different competitions since the scoring resets each time, but within the same competition, it's more of a problem. 

I looked at the 10 Games events in which all athletes competed, which gave me a total of 45 unique combinations for men and 45 combinations for women. Of those combinations, only 8 had correlations greater than 50% and only 3 had correlations greater than 70%. Not surprisingly, the 2K row and the half-marathon row were highly correlated for both men and women (54% for men and 81% for women). Also, the Sprint Chipper and the C&J Ladder were strongly correlated (70% for men and 54% for women), likely because they both had a major emphasis on heavy Olympic lifting. One surprise was that the burden run and the 2K row were 79% correlated for women, but I think that may have been somewhat of a fluke, considering the correlation was just 31% for men.

In the end, most events appeared to test pretty distinct aspects of fitness, which is a good sign.

How We Can Do Better: Fans love the heavy movements, but I'd suggest supplementing those with some more moderate weights as well. CrossFitters can relate to someone crushing a workout even if the weight it not enormous (those Open WOD demos weren't bad to watch, were they?) Also, let's test running earlier in the season.
Credit Where Credit is Due: We saw events where even Rich Froning and Sam Briggs found themselves near the bottom, which tells me we are really testing a wide range of skills. And actually, I liked limiting the Games to 12 events (instead of 15 last year), because in my opinion that was sufficient and we didn't wind up double-counting too many areas.


5. Make the event enjoyable for the spectators

Unfortunately I don't have any data to back this up, but in my opinion, this is the area that I think has improved the most in recent years. I think a nice touch at the Games is that in multi-round workouts, each round is performed at a different point on the stage. This really helps the audience follow the action and builds the drama as you see athletes progress through the workout.

Making all the events watchable was also nice after Pendleton 1, Pendleton 2 and the Obstacle Course were unavailable last season. The burden run had many of the same qualities as an off-road event, but it was all done on site and finished up in the soccer stadium.

However, as nice as it is to use the soccer stadium to allow more spectators, the vibe at those events is considerably more subdued. Perhaps HQ will be able to find a way to improve this in the future, but it seems that this sport isn't quite as conducive to viewing from such a distance. By contrast, the intensity in the night events in the tennis stadium is fantastic.

How We Can Do Better: Figure out a way to make things a bit more exciting in the stadium. It won't be easy, but there's no denying that things weren't quite as intense when the workouts were held there.
Credit Where Credit is Due: The Games are truly becoming more of a spectator sport. Even the uninitiated can see the action unfold and understand and appreciate what's going on. And although I mentioned it above, the improvements in judging have helped the spectator experience.


*I decided to break up "wall burpees" into burpees and wall climb-overs. Each were worth 1/6 of the value of that workout (snatch was 1/3 and weighted GHD sit-up was 1/3). This was updated on 8/22/2013.

Friday, April 19, 2013

A Look Back at the 2013 Open: Part II


Note: For details on the dataset I am using, including which athletes are and are not included, please see the introduction to Part I.
 
For Part II of my look back at the 2013 Open, let's transition into looking to some results from this year's Open, while still keeping the programming in mind. As I mentioned in the intro of Part I, one way to judge the effectiveness of a workout (for competition purposes) is to see how well it correlates with success in a variety of other events. In other words, an event is "good" if the best athletes tend to finish near the top. Although we don't yet have regional results to add to the mix, we can still look at how well each event predicted success in the other four Open events. By this metric, event 13.4 was probably the best test this season. For the entire field, a male athlete's rank on 13.4 was 91% correlated with his rank on all other events combined (for women, 90%). All other events ranged between 83-86% for men and 80-88% for women. 

If we limit the field to the top 1,000 finishers for men and top 500 finishers for women (roughly 2% of the final field for each), re-rank everyone, then look at the correlations, they do drop off. This makes sense, because it is not inconceivable that a top athlete might fall toward the bottom of that elite group on one event, but a top athlete falling anywhere beyond the top 5-10% across the entire field is unlikely.  Looking at the events across this elite group, 13.5 had the highest correlation for men at 39%, while 13.4 was still highest among women at 53%. Among this elite group, 13.2 proved to be the weakest test for both men and women (23% for each), which is not particularly surprising to me. The difference between top scores on this workout often came down to tiny fractions of a second on each rep, not to mention that the was widespread variation in judging standards on this one (which has been discussed on the internet ad nauseum already).

To put these numbers into some context, I looked at the correlations for the five events last year for the men's field. Across the entire field, events 12.3, 12.4 and 12.5 had correlations between 81-88%, but the two single-modality workouts, 12.1 and 12.2, only had correlations of 69% and 63% respectively. Limiting the field to the top 500 (roughly 2%), the correlations for 12.1 and 12.2 were just 15% and 12%. Not to beat a dead horse, but the multi-movement events just give us more information about the athletes, which is key in an competition as short as the Open.

Visually, these scatter plots help illustrate the idea. Each chart shows a random sampling of about 500 athletes from across the entire field (because showing them all would be too much for Excel, which is why I need to improve my R skills). On each chart, the location on the x-axis represents the athletes's rank on that event, and the location on the y-axis represents that athlete's combined rank on all other events. Notice how much more of a clear relationship we have in 13.4 as compared with 12.1. Virtually no one did amazingly well on 13.4 without also performing solidly across all five workouts.




Moving on, with the 2013 Open web site being set-up in conjunction with last year's site, it was much easier to track which competitors were returning and which were new. With that information available, I was curious to see how the returning athletes fared as compared to the first-timers. Would they have more of an advantage on certain events. 13.3 perhaps? Let's see.*


Simply put, the returning athletes did fare better than the newcomers (lower percentiles in this case mean a better ranking). However, the difference was not markedly different on any one event. I anticipated that returners would have an even bigger advantage on 13.3, given that it was in last year's Open, but this does not appear to be the case. Interestingly, among the top 1,000 male finishers, new athletes actually fared slightly better than returning athletes on 13.2. This is the only event where this was true, and it was actually not true for women.

What would be ideal here is to have the age for each competitor so I could control for any differences in the age mix between the three categories of athletes, but alas I do not have that information at the moment. My hunch is that the age mix is likely pretty similar between the three groups, however.

Since 13.3 was an exact repeat of 12.4, it should provide some good insight into how the community is progressing over time. In order to better compare performances on most events, I like to adjust the scores to be in terms of stations. So in this case, 150 wall balls is 1.0 stations, 90 double-unders is 1.0 stations and 30 muscle-ups is 1.0 stations. For this event in particular, this isn't a perfect solution, because 90 double-unders is a considerably shorter station than the other two, but I think this still gives us an improvement. We know that 1 wall ball is much less difficult than 1 muscle-up, so we try to reflect that.

After making that adjustment, I compared the distribution of scores across all competitors on 13.3 vs. the distribution of scores across all competitors on 12.4. The charts for men and women are similar, but since I've been showing men first up to this point, let's show the graph of the ladies here. The chart below is basically made up of two histograms, which I've shown as line because it makes it easier to compare their shapes. I've also omitted scores below 0.75 rounds (113 reps) because the graph tails off quickly and then has a spike at the very bottom due to all the scores of 1.


I found this truly remarkable. Despite doubling the size of the field this year, the scores across the entire community were distributed in nearly identical fashion. Additionally, the percentage of women completing a muscle-up inched up from 9.9% to 10.2%, and the percentage of women completing a muscle-up of those that reached the muscle-ups ticked up from 27.0% to 28.6% (these numbers are not easy to discern from the graph). For the men, the percentage completing a muscle-up actually dropped from 37.1% to 35.6% and percentage completing a muscle-up of those that reached the muscle-ups stayed flat at 73.6%. So essentially, no real change from last year.

Does this mean that CrossFit somehow doesn't work? Are athletes not improving at all with a year's worth of training? No. What it means is that the newer athletes are coming in and putting up scores at the lower end of the spectrum, offsetting the athletes who competed last year and are improving. How do we know this? Well, let's look at the same chart as above, but this time only with athletes who finished all five events in 2012 and 2013.


You can clearly see that the distribution of scores on 13.3 is heavier at the higher end than the distribution for 12.4. For instance, you can see that about 15% of athletes scored between 2.0-2.25 (240-247) on 13.3, while only about 8% did so on 12.4. Further, the percentage of women completing a muscle-up jumped up from 12.6% to 23.9%, and the percentage of women completing a muscle-up of those that reached the muscle-ups went up from 30.2% to 40.9%. For the men, the percentage completing a muscle-up skyrocketed from from 49.8% to 70.9% and percentage completing a muscle-up of those that reached the muscle-ups moved from 77.8% to 85.3%. Clearly, the athletes who came back this year worked on wall balls, double-unders and muscle-ups, and it paid off on 13.3.

Finally, let's wrap things up by taking a look at how many athletes stuck with it through all five weeks of the Open. As has been noted in the past, scores from athletes who only compete in the first event or two are not removed from the standings for the weeks in which they did compete. Although strangely, athletes who skip week 1 but compete later do not count at all. I would argue this is not exactly fair, as the first events essentially get weighted more because the size of the field is larger. My analysis, however, did not show that this would have had a significant effect on the final rank of competitors.

Still, it is interesting to look at when and how rapidly athletes drop out of the field. Here are stacked bar charts for the men and women in this year's field. If an athlete returns to the field after dropping out, that score after returning does not count in my analysis.


You can see only thing out of the ordinary there appears to be the number of women finishing the first 3 events. My hunch is that this is due to the large number of women who could not complete a 95-lb. clean and jerk. In fact, let's look back at how the field tailed off for men and women the past two years, from a slightly different perspective.


Again, the only drop that looks out of line is the women going from 13.3 to 13.4. Between those events, the field shrunk by 19%. In the past two years, the next-largest drop was 13% (men between 13.4 and 13.5), and all others were between 9% and 11%. I would have liked to look back at 2011 to see if a similar percentage dropped off from 11.2 to 11.3 (the heavy clean and jerk), but the 2011 Games site is considerably more challenging to handle. If anyone has an answer there, I'd certainly be curious.

Anyhow, that's it for now. There may be more topics to re-visit from this year's Open, but I think it's time to focus our attention to Regionals (even for the vast majority of us who won't be competing). See you all in a few weeks when the events are announced!

*"Returning incomplete competitors" refers to those who started last year's Open but did not complete all 5 events.