Follow me on Twitter!


Showing posts with label Curiosities. Show all posts
Showing posts with label Curiosities. Show all posts

Tuesday, April 1, 2014

Can Mid-Week Projections Work?

Two weeks ago, I proposed a method to project an athlete's overall ranking before score submissions had closed for the week. To me, it made sense on paper, but it was admittedly untested. So I put out a request for help on testing it in week 4, and thanks to Andrew Havko (among others), I was able to make that happen.

So can it work? It appears that it can. That's not to say the projections are 100% accurate, and they are far from precise very early each week. But I think it's clear that the projections can give an athlete a good sense of where they would likely finish the week if they stick with their current score, which is something that is nearly impossible currently.

I tested these projections at three points during week 4: Friday 8 a.m., Saturday 5:30 p.m. and Sunday 3:30 a.m. (all EDT). The method requires one key assumption, which is the percentage of athletes who will drop off from the prior week, and for this I used 10%. Certainly this would need a bit more careful thought if it were to be implemented by HQ.

For each athlete, I projected their overall worldwide ranking at each of these times. For athletes whose score did not change by the end of the week, I compared my projection to their ultimate ranking. In total, the error of my projections were as follows:
  • Friday 8 a.m. (<1% of field reporting) - 9,575 mean absolute error*, 9,404 mean error
  • Saturday 5:30 p.m. (16% of field reporting) - 1,003 mean absolute error*, -787 mean error
  • Sunday 3:30 a.m (21% of field reporting) - 1,454 mean absolute error*, -1,362 mean error
Interestingly, the projections (at least using this first basic method) got slightly worse overall from Saturday to Sunday. The reason is that the distribution of scores submitted by Saturday 5:30 p.m. was more similar to the ultimate distribution than on Sunday. What I found was that, in general, the scores submitted very early on during the week are well above average, and the quality slowly declines throughout the week.  That is until Monday evening, when a slew of athletes replace their first score with a second improved submission. It turned out in this case that Saturday afternoon was a pretty accurate indication of how the current week's scores will turn out.

However, let's look a little more closely at the errors. Although an error of 1,003 (our best mean absolute error) is pretty small for an athlete finishing, say, 40,000th, it would be a very large error for an athlete finishing 2,000th. Thankfully, the size of the errors generally increased as the ranking increased. Below is a chart showing the percentage error for athletes across the spectrum of rankings, using our Saturday afternoon projections.


So you see that generally, we never really stray further than 3% error at any point. That's not too bad when you consider that there's currently no way to get even a good ballpark estimate until at least mid-day Monday.

Still, maybe we can do better. What if we had actually used the perfect assumption (8% in this case) for the percentage of athletes who would drop off from the prior week?  Well, in total, we improve for our Saturday and Sunday projections, with the mean absolute error going down to 338 for Saturday and 581 on Sunday. Interestingly, though, in this particular case it doesn't necessarily improve the projections across the board for Saturday and Sunday. Below is the same chart as above, but with the perfect assumption for attrition.


Although our error gets a little worse near the top, once we get near the middle of the pack, these projections are nearly spot-on. And even near the top, a 5% error isn't that bad - that's like these projections putting Josh Bridges at 100th overall, whereas he actually finishes 105th.

One way we can theoretically adjust to get even closer is to make an adjustment for the skill level of the atheletes who have submitted scores at a given point. This could involve looking at the average ranking of the athletes from their prior week's scores and comparing that to what we'd expect by week's end. The trouble is, it's challenging to know what the level will be at week's end. You might expect that the field would average out to be at the 50th percentile in prior weeks, but that wasn't actually the case here. The average athlete submitting a score for 14.4 was actually about the 48th percentile in prior weeks, which is due to the fact that the athletes dropping out after 14.3 were generally from the bottom of the pack.

My point is that while such an adjustment is possible, it might not be practical. And considering the projections even with my base 10% attrition assumption weren't too bad, I don't think further adjustments are necessary, beyond refining that attrition assumption to make it as accurate as we can.

Finally, while I think this method would produce reasonable results if implemented by HQ next year, there are some caveats about the testing done here:

  • I've only done testing for one week. There may be more (or less) error if we made these projections in week 2 or week 5.
  • I'm almost certain that the percentage error would increase a bit if we do this for each region. The sample size is much smaller, which means that even if the same principles apply, we're likely to see more variability. For one thing, it's going to take longer each week before the projections are even remotely meaningful, since many regions had less than 100 entries until late each Friday afternoon.
  • I only tested this for the men's field. I don't see any reason why the results would be much different for women, aside from the field being smaller, which would likely increase our percentage error a bit.
All that being said, I feel that implementing this method would provide a realistic glimpse into where an athlete will wind up. As long as athletes understand that this is merely an estimate, the information provided can be quite useful. 

Would this revolutionize the sport? Of course not. But I think it would be yet another improvement to the athlete experience as the largest stage of our sport continues to grow.


*Mean absolute error is the average of our errors, if we ignore the direction of the error. So if we are off by -500 for one athlete and +500 for another, the mean absolute error is 500 but the mean error is 0.

Monday, March 17, 2014

A Method to Project Overall Open Rankings Mid-Week

One quirk about the Open leaderboard is that while a workout is open for submissions, the overall rankings are basically useless. The rankings for the current week's workout are obviously understated, but as I explained in my previous post, you can at least get a decent sense of where the score will end up by looking at the percentile rank at any point in time. However, with the overall rankings, they are screwed up because the most recent week's rank is so understated in relation to the prior weeks' scores. For instance, if an athlete who was in 300th place in each of the first two weeks but posts the best score in the world early in week 3, he will still appear behind an athlete who finished 290th in each of the first two weeks but is currently 10th of 100 entries in week 3. But we know that by week's end, there will be much more separation between the athletes in their week 3 ranks, which will place the first athlete well in front.

So is there a way we can get at an accurate projection of an athlete's overall ranking mid-week? I think we can, but not without a little bit of work.

The idea is this: since we can reasonably project the ending percentile ranking for the current week's workout, we should be able to reasonably project the ending rank, if we make an assumption about how many athletes will complete the workout. If we can get that projection for any particular athlete, we should be able to do that for all athletes who have completed the workout. At that point, we can re-rank those athletes based on projected total points. Using that, we can basically "scale up" those ranking based on how many athletes we anticipate will complete the current week's workout.

More specifically, here is the process I am proposing:
  • Compute each athlete's percentile ranking for the current week (either overall or in the region) based on the athletes who have currently submitted scores
  • Based on the number of athletes in contention at the end of the prior week, reduce that by some factor (say 10%, which is near the historical average) to get an estimate of the number of athletes who will remain at the end of the current week
  • Multiply the athlete's percentile ranking by the estimated number of athletes who will remain at the end of the current week to get the projected rank for the current workout
  • Use these projected ranks to get a projected overall point total at the end of the current week
  • Re-rank the athletes who have submitted scores based on the projected point totals
  • Convert the projected rankings to a percentile rank based on the number of athletes who have currently submitted
  • Multiply this percentile by our earlier estimate about how many athletes will remain at the end of the current week. This will give you each athlete's projected overall rank at the end of the week.
To accomplish this, all we would need a snapshot of the full leaderboard at a given point in time. I do not think it is possible to accomplish this even for a single athlete without making the calculations for all athletes. However, with the right computing power, it would be a relatively painless calculation to generate the projected overall rankings. Obviously HQ would be in the best position to perform these calculations, but I think it is conceivable that someone on the outside could do this as well.

This is all theoretical at the moment - a decent amount of testing would be necessary to make sure this process actually produces reasonable projections. Still, I think the concept is something that could be used to improve the Open experience for all of us.

Note: If anyone out there has the resources and the know-how to get a hold of the leaderboard mid-week and get it into Excel or .csv, I'd be very interested to test this out. If so, post to comments or email me directly (anders@alumni.wfu.edu).

Friday, December 20, 2013

Are CrossFitters Specializing?

For the past seven years, the CrossFit Games have sought to find the fittest all-around athletes in the world. One common theme from CrossFit HQ is that the Games seek to "punish the specialist" and "reward the generalist." However, one can't help but notice the increased attention that certain aspects of fitness seem to garner in the CrossFit community compared to others. In CrossFit media, the emphasis on lifting, and Olympic lifting in particular, seems to be disproportionate to many other areas of fitness. When was the last time you saw a video or even a note on the Games site mentioning a Games athlete hitting a new PR in his 5K run? Yet it seems like it hasn't been 24 hours since we've seen a video posted of another athlete hitting a big snatch or clean and jerk.

OPT noted this phenomenon on his blog about a year ago:

"Media recently for the sport has put an emphasis on strength development in spite of promoting true “balance” in fitness and the general components of fitness.  A sport where now the elite can qualify for the American open weightlifting championships but cannot qualify for a state-level high school cross country meet."

Although I won't seek to prove it in today's analysis, I think there is little doubt that the big lifts tend to get a lot more attention than general metabolic conditioning in CrossFit media. However, the question I will attempt to answer is whether the sport itself has gotten out of balance.

From the programming perspective, I showed in my recent post "History Lesson: An Objective, Analytical Look at the Evolution of the CrossFit Games" that while the metcons have become heavier and heavier over time, the overall balance of lifting and conditioning has not changed drastically in the past 7 years. In addition, there is roughly the same amount of emphasis on Olympic lifting now as there has been throughout Games history. In fact, there has actually been a shift away from the powerlifting-style movements like the deadlift and back squat. Running is and has always been the most common movement at the CrossFit Games.

However, there is a legitimate question about the intense focus on Olympic lifting and the lack of focus on running at the Regional level. And there is also no doubt that the CrossFit athletes have been getting more and more proficient at the Olympic lifts, as evidenced by the rising numbers in the 1-rep max events at the Games each year.

So let's ignore the programming for now and focus on the actual strengths and weaknesses of the athletes in our community. Before I do this, I want to re-visit the comment from OPT above. I love OPT and have probably watched every CrossFit.com video of his over the past few years, but I think the commentary that "the elite can qualify for the American open weightlifting championships but cannot qualify for a state-level high school cross country meet" is a bit misleading on a couple of levels:
  1. The American Open is not that competitive on a global scale. To qualify in the 85 kg weight class, you need a 266 KG total (http://0205632.netsolhost.com/2013NationalEventsQualifyingTotals.pdf). Yet at the 2012 Olympics, the 16th place finisher had a total of 315, or 18% higher. On the other hand, to qualify for the highest-level state cross country meet in Ohio (a competitive state where I used to cover sports), you need a time around 17:00. Considering these meets are run on rugged terrain rather than on a track, this time isn't that far behind the Olympic 5,000 meter times (13:52 was 15th place in the 2012 Olympic final). So I think it could be argued that qualifying for those two events are actually relatively comparable as far as difficulty.
  2. There are weight classes in Olympic weightlifting, yet there are not in CrossFit or cross-country. The top CrossFitters are not even in the same stratosphere as the lifters in the 105KG+ weight class. On the flip side, among runners above 200 lbs., I have to believe someone like Garrett Fisher would be considered elite.
But let's look at the numbers throughout our community. To do this analysis, I used the 2013 Open data, which was generously pulled and cleaned for me by Michael Girdley (girdley.com). This dataset has all the numerical information provided by athletes who competed in the Open (it does not include answers to the questions about diet, how long you've done CrossFit, etc.). Based on this self-reported data, I believe we can understand how CrossFit athletes from top-to-bottom compare to the world's best in a variety of lifts, running events and metcons.

To perform this analysis, I first limited the data to athletes under 40 who completed all five events (approximately 39,000 men and 23,000 women). Then I re-ranked all the athletes based on their rank across all 5 Open events and grouped them into 20 buckets based on this rank. Within each bucket, I took the average for each of the self-reported scores (Fran, Helen, Grace, Filthy 50, FGB, 400 meter run, 5K run, clean and jerk, snatch, deadlift, back squat, max pull-ups). For the timed events, I converted these to a pace (rounds per second, for instance, or meters per second). Then, I pulled in world records for each and compared the CrossFit community against those world records.

The charts below shows how the community compares to the world records*. For the lifting events, these are the world records without regard to weight class, since CrossFit does not have weight classes. To reduce clutter, I have grouped all metcons together, all runs together and all lifts together (pull-ups stayed in their own category).





On both charts, we can see that the community is generally closest to the world record when it comes to running events and (not surprisingly) metcons. When compared to the world record in the lifts, it becomes obvious how far behind the CrossFit world still is. Proud of that 200-lb. snatch? Congratulations, you are slightly below half the world record. (Note: I am proud of my 200-lb. snatch, and it took me 5 years to finally get there).

But let's also look at the Games athletes in particular and see how they stack up. Here is a table showing each event and how the average Games athlete stacks up compared to the world record.


Not shockingly, the Games athletes are near the top to the world record in the metcons (since generally the world record comes from this field), but again, note that they are much closer to the world records in the 5K run and the 400 meter run than they are in the Olympic lifts. And look at that back squat - not even close! (and before you ask, this is comparing against the raw world record in the back squat, which appears to be 450 KG as best I could tell)

It is interesting to note that the elite CrossFit men are closer to the Olympic lifting world records than the elite women, yet they are further from the 5K run record. This could have something to do with the background that many of the athletes had prior to CrossFit, but that's purely a guess at this point.

Another way to look at this is to understand where the Games athletes excel furthest beyond the rest of the community, and in particular, where they excel furthest beyond the rest of the Regional field. For both men and women, here is a look at how the top 5% of Open finishers (roughly the Regional field) compared to the Games athletes.


Here I think we start to see something interesting. While the Games athletes do not appear to be any further from the world record in the 5K run than they are in the Olympic lifts, they aren't that much better than the rest of the regional field when it comes to the running events. They also aren't that much better when it comes to the powerlifting movements. I think you could attribute at least partially to programming at the Regionals: we simply aren't testing much for running or powerlifting, so the athletes making the Games aren't necessarily that much better than the rest of the field in those areas.

On the flip side, look where the Games athletes do exceed their peers by a greater amount: the Olympic lift, the short metcons and pull-ups. It seems that explosive power and conditioning (over a relatively short time frame) are what tend to separate the Games athletes from the rest of the Regional field.

One last way to look at this is to see the gap between the Regional athletes and the median Open athlete**, which is defined as the athletes finishing in 45th-55th percentile in the Open among people under 40 who completed all 5 events. These median Open athletes are still generally fit individuals, they just aren't quite at the Regional level.


This table looks a lot like the prior one, meaning that what separates the Regional athletes from the average Open athletes is a lot like what separates the Games athletes from the Regional athletes.

Based on the analysis here, I believe that CrossFit athletes in general aren't bad runners or particularly tremendous lifters. However, the elite CrossFit athletes are significantly better lifters than the rest of the community, and yet they are not drastically better runners than the rest of the community. From this perspective, we do see a little bit of the bias that OPT was writing about. But overall, I don't think the specialization issue is as much of a concern as some might think.

Update 12/20: [I'd like to note that I don't believe that achieving 70% of the world record in the snatch is exactly as challenging as achieving 70% of the world record pace in a 400 meter run. However, the fact that CrossFit Games athletes are so much closer to the world record in the running events than they are in the lifts indicates to me that these athletes should not be considered specialists in the Olympic lifts who simply neglect running. One way to quantify this, which I'm hoping to look into more, is to put things in terms of standard deviations. I have looked at this for the snatch, clean & jerk, 5K run and 400 meter run for men, however. Using the standard deviation based on the same sample of Open athletes under 40, the Games athletes are approximately 6.0 standard deviations below the world record in the lifts but only 4.3 standard deviations away in the 5K run and 2.0 standard deviations away in the 400 meter run. This isn't a perfect method either, but again, it supports the idea that Games athletes aren't totally specializing in the Olympic lifts while neglecting their running.

However, I do see the same pattern as in the main body of my post when comparing Games athletes to the rest of the CrossFit field. Games athletes are only about 1.1 standard deviations better than the median in the 400 meter run, 0.7 standard deviations better in the 5K run, but approximately 2.5 standard deviations better in the Olympic lifts. So it seems that the same conclusions generally hold when doing the analysis this way.]


Update 12/21: [As a follow-up to the previous update, I looked at where the average Games athlete would fall in the spectrum of all Open athletes in each of the self-reported metrics. This was more difficult than it might seem because of the tremendous selection bias in the data (only about 20% of the men's field reported a 400m time, for instance, but about 50% reported a deadlift max). I tried to account for this by creating a "weighted" distribution, where each 5% bucket was only worth the same number of total athletes, regardless of how many missing values they had. After doing this, I found that the average male Games athlete is in the top 2% in the clean and jerk and snatch, and they were at least the top 6% in all other lifts or metcons. However, for the 400 meter sprint, they were only in the top 15%, and in the 5K run, they were only in the top 25%.

Note that the selection bias still can't be totally accounted for. It's probably fair to assume that the non-responders in general had worse scores than those that did respond, so maybe the Games athletes are actually even better than they appear here. However, it is definitely striking that the Games athletes are not that far beyond their peers in the runs, particularly the 5K. Still, it doesn't necessarily say they are bad runners, as you could argue that CrossFitters in general are good runners and therefore the Games athletes are still pretty good. It is clear, however, that Games athletes are outdistancing their peers substantially in the Olympic lifts, even though they are generally still well short of elite status. 

I think a lot of this has to do with the background of many CrossFitters. Many, many people ran to stay in shape prior to finding CrossFit, but relatively few Olympic lifted. To really decide if you feel the sport of CrossFit has gotten too specialized, I think all of the preceding analysis has to be taken in together, including how far CrossFitters are from the world records as well as how the Games athletes compare to the rest of the field. I'm not sure there is really a clear-cut answer.]

Update 12/28: [Quick one here. I did the same analysis for the women that I did for the men on 12/21, and I found that the women's Games were slightly more dominant across the board. The average Games athlete would be in the top 3% for all lifts and metcons, the top 9% for the 5K run and the top 15% for the 400 meter sprint. Interesting that they were comparatively better than the men in the 5K, although actually about the same in the 400 meter sprint. I'm not sure I really have a good hypothesis for this at the moment.

Also, worth noting is that I looked into the response rates for each metric, and found that for women, the rate was between 7% and 12% for all runs and metcons, except Fran, which was 17%. The response rate was between 32% and 38% for all the lifts and 16% for max pull-ups.

For men, the rates were between 13-23% for all metcons, except Fran, which was 33%. The response rate was between 44% and 51% for all the lifts and 30% for max pull-ups.

This does indicate that there is a selection bias issue that has to be considered, but it's not as if it ONLY applies to the runs. Basically all the metcons and the runs had very low response rates, but the lifts had much higher response rates.]

*Here are the world records I used in this analysis, based on a combination of web research and self-reported PRs from the database: 
Fran - 2:00 (men), 2:07 (women)
Helen - 6:13, 7:20
Grace - 1:14, 1:17
Filthy 50 - 14:05, 16:13
FGB - 520, 460
400 meters - :44, :50
5,000 meters - 12:37, 14:11
Clean & Jerk - 263 KG, 190 KG
Snatch - 214 KG, 151 KG
Deadlift - 461 KG, 264 KG
Back squat (raw) - 450 KG, 280 KG
Max pull-ups - 106, 80

**There is a significant amount of selection bias in these self-reported numbers, which is why I used the bucketing approach to account for it. In general, the people reporting their numbers for each lift/run/metcon are better at those lifts/runs/metcons than those who leave them blank. Also, for many of the metcons, less experienced athletes may not even have a PR. As an example of this bias, if you take a straight average of the clean and jerk across all women under 40 finishing all 5 events, it's about 134 pounds. But if you group the field by the 5% buckets as I have, take the average in each bucket, then average across all buckets, you get an average of 126 pounds, which I believe is more representative of the "true" average.

Sunday, June 30, 2013

Which Have Been the "Best" Events This Season?

If you've been reading my blog for any length of time, you know that one of the ways I like to evaluate the effectiveness* of a CrossFit event is by looking at how well the results from that event correlate to results from a variety of other events. I laid out the theory behind this in my post from last year titled "Are certain events 'better' than others?", but I'll recap it here:
  • In a competition setting, what we are trying to do is learn as much about each athlete's overall fitness level as possible.
  • We only have a limited number of events to do this, particularly in the Open and at Regionals. Therefore, we need to maximize the information we get from each event.
  • If athletes who score well on a particular event tend to score well across the board, then that event is probably a good indicator of overall fitness. Conversely, if the results from that event don't correlate at all with results in other events, then maybe that particular event did not really tell us much.
Overall, this year's regionals and Open were set up well for me to do this type of analysis. I did this same analysis after last year's regionals, but due to the cuts after regional event 5, I was left with only about 250-300 athletes of each gender who completed all the Open and Regional workouts, and this limited the analysis to only the very elite athletes. I also did this analysis after the Open this year, which gave me a huge sample of athletes, but with only 5 events, I didn't really have a "wide variety" of events to evaluate. 

This year, because we did not have any cuts at Regionals, I was left with 673 men and 512 women who have completed all 12 events this season. Although these are all still very solid athletes, I got a lot more of the borderline regional competitors in the mix than I did last year. Remember, a "good" event for the Games might not make a "good" event for a competition within your own box. So keep in mind with this analysis that we are evaluating these events based on how well they predict overall fitness for Regional competitors.

The methodology for performing this analysis this year was the following:
  • For athletes who completed all events this year, compile their results (not just their ranking) for each of those events. The Regionals results I used are those that have been adjusted to account for the week in which week each athlete competed. See my previous post for more info on those adjustments.
  • Rerank the entire field in each of those events.
  • For each event, calculate the sum of each athlete's ranks in all other events.
  • Calculate the Pearson correlation (referred to from here on out as simply "correlation") between the ranks in each event and the sum of ranks in all other events. Higher correlations indicate "better" events.
Below are the results for both men and women.















The pattern that emerges is one I've noticed pretty much across the board since I've been doing this type of analysis: events with more movements tend to be better tests of fitness. This makes sense intuitively, since they test more things, by definition. That doesn't mean we shouldn't have single-modality events in competition, it just means they probably should be used sparingly and only for movements that are deemed very important.

You may notice that Open Event 2, which had three movements, is bucking this trend by falling quite close to the bottom. The concerns with this event were pretty well documented during the Open: judging was very difficult, the weights were extremely light for top competitors, and the option of step-ups was utilized by a lot more athletes than was probably expected. So simply because you have 3 or 4 movements in an event doesn't make it a great one.

To get a visual interpretation of the concept I'm getting at here, below are scatter plots for some of the best and worst events from this year. On each graph, the x-axis represents each athlete's rank on that event and the y-axis represents each athlete's combined rank on all other events.




It should be fairly clear that for the first two graphs, there is a clear relationship between the x- and y-axis. Athletes who did well on these events generally did well across the board. On the third graph, the points are much more scattered, indicating that there were plenty of athletes who did well on Open Event 2 who didn't fare well across the board, and vice versa.

Although I consider the results using the entire Regional field to be the most useful, I also tested performed this analysis on three other subsets of the regional field:
  1. Games athletes only
  2. Top 292 men and top 258 women (same number as I had in my analysis last year)
  3. Random sample of 20% of the entire regional field

What I was interested in was how volatile my results were. For instance, is Regional Event 1 really better than Regional Event 4 for women (76% vs. 74%)? Would that hold up if I changed the group of athletes a bit?

In general, the events near the very top and the very bottom stayed in that vicinity, with a couple of exceptions. Here are the main takeaways:
  • For men, Regional Event 4 was in the top 3 across all the samples and Open Event 1 was in the top 5 across all the samples. For women, Regional Event 4 was in the top 4 across all the samples, Regional Event 1 was in the top 5 across all the samples, Open Event 1 was in the top 5 across all the samples and Open Event 4 was also in the top 6 across all the samples.
  • Considering HQ is programming the same events for both men and women, I would conclude that Regional Event 4 ("The 100s") and Open Event 1 (AMRAP 17 of burpees and snatch) were the best events this season. In one of my 2013 Open recap posts, I noted that 13.4 was generally the best event of the Open. I still feel that it was a very good event for the entire Open field, but it wasn't quite as strong when we look at just these stronger athletes. 
  • Across both men and women, Open Event 2, Regional Event 5 and Regional Event 3 were each in the bottom 4 in all but one sample. I would conclude that these were generally the three weakest events this season. 
    • I mentioned some issues with Open Event 2 above.
    • For Regional Event 5, I think the issue is that we saw a lot of athletes near the bottom of the field do well on this simply because they could deadlift a house. If you could handle the deadlifts easily, you could generally do well even if you weren't particularly great at box jumps or had sub-par aerobic capacity.
    • For Regional Event 3, I think the issue was that the burpees did not really factor into this much, making it basically just a muscle-up test. As far as single-modalities go, this wasn't too bad of an event. But personally, I think this event and the overhead squat ladder (a true single-modality) should have been worth only 50% of the other events.  
  • Two of the events featuring box jumps turned out to be relatively modest tests of fitness. Throw in all the complaints about achilles problems we've seen popping up recently, and I think HQ may want to look into adjusting how they program box jumps. I think box jumps are a good test of fitness in general, but I'd personally love to see us go to box jump-overs (onto and over the box with an option to jump straight over) in the future.
  • With the exception of Open Event 2 and Regional Event 5, I think the rest of the events were generally solid. As I mentioned above, I might consider adjusting the point value for a couple of the other ones.

I also did one final analysis, primarily out of curiosity. Using the entire Regional field, I looked at the correlation between each pair of events. Some of the interesting findings:
  • In general, the most highly correlated pair was Regional Event 4/Regional Event 1 (71% for women and 67% for men). This is somewhat surprising given how the time domains were completely different, but both involved pull-ups and a light thruster-type movement (bar thrusters or wall-balls).
  • The two muscle-up workouts (Regional Event 3 and Open Event 3) were highly correlated (68% for women and 55% for men).
  • Regional Event 5 and Regional Event 7 were highly correlated (56% for women and 68% for men). Both were extremely heavy.
  • The least-correlated pair was Regional Event 3/Regional Event 5 (17% for women and 19% for men). Shouldn't be a surprise considering one was bodyweight only and one involved extremely heavy deadlifts. The pair of Regional Event 2/Regional Event 3 were also not very correlated (38% for women and 22% for men). Remember those occurred within 2 minutes of each other.
That's about it for today. This is always one of my favorite analyses to work on, but with the Games fast approaching, I suppose it's about time to start tackling the tough questions and making some predictions. Will Froning three-peat? (Probably) Who will emerge on the women's side with Annie out? (It's wide-open) Will the first event of the Games take more or less than 4 hours? (God, I hope so) What bizarre contraption will Rogue unveil this year? (Potentially a flying bicycle, similar to the one in E.T.) Who will wear the shortest shorts this season? (Stacie Tovar still the champ until proven otherwise)

Anyway, until next time, good luck with your training!


*I am referring to the effectiveness of this event as it relates to competition. In other words, is this event a good test of fitness. This does not necessarily mean the event is good or bad for training purposes. For instance, I feel that 13.2 was not a good workout for testing (due to a lot of factors, like how difficult it was to judge and how light it was for the top competitors), but in training, I think it would be a good workout for building aerobic capacity (and it definitely left me hurting).

Monday, December 10, 2012

Quick Hits: Women's Scaling and The Outlaw Open

In stark contrast to my last post, I just wanted to touch on a few topics briefly today without getting too involved. These were just a couple of interesting topics that I got into this past weekend.


Women's Scaling:

CrossFit HQ has asserted in the past that there are no "prescribed" women's scaling options for their workouts. According to HQ, each workout should be scaled to the particular athlete's abilities. I'm not going to argue with that sentiment here, at least as it applies to training. But the fact is that HQ does scale the weights differently for men and women during competitions, and most, if not all, local CrossFit competitions do the same. But are they scaled fairly?

Counting the Outlaw Open (Dec. 1-2), I've now analyzed seven different competitions from the past two years (the other six are the Open, Regionals and Games from 2011-2012). What I wanted to look at now was the average relative weight on metcons and the load-based emphasis on lifting (LBEL) in each competition for both men and women. What I found was this:

For metcons, the average relative weight for women has been between 61-71% of the men's load. In total, the LBEL (which includes all events) has been between 61-69%. That's not a huge spread, but it could mean the difference between programming cleans at 95 lbs. and 110 lbs.

To see if there was an "optimal" relativity between men's and women's weights for CrossFit competitions, I decided to look into the (few) non-metcon lifting events we've had in those competitions. In those seven competitions, there have been three max snatches, two max cleans, one max thruster, one max bench and one 5-rep max deadlift. Taking the average of the field in each of those events, the women's results were between 51-65% of the men's. The 51% was on the bench at the Outlaw Open - my theory on the bench being so much lower is that a lot of emphasis is placed on the bench in many men's sports, including football (just a theory though). Other than bench, all others were between 58-65%. Considering that the average female competitor at the 2012 Games weighed a shade under 75% of the average male competitor, I'd say those are some damn impressive strength numbers.

If those maxes are a decent indication - and I think they are - then the programming should generally be looking to keep the women's weights at around 60-62% of the men's. And in fact, things appear to be moving toward that. In 2011, the women's loads in the Open, Regionals and Games were between 66-71% of the men's (using either LBEL or average metcon weight); in 2012, they were between 61% and 68%. At the Outlaw Open, the LBEL was 62% of the men's and the average metcon load was 65%.

This all assumes that the goal is to make the women's events just as challenging for their field as the men's events are to their field. For the most part, bodyweight movements are not scaled at all for women, which helps explains why in the 2012 Games, the average women's time on metcons was still about 20% slower than the men's. But perhaps that's not a problem. I think a lot of the draw for women's competition is seeing these women doing virtually the same thing as the fittest men in the world. I'd venture to say the top women are about as proficient at muscle-ups as the top men were just 3 or 4 years ago, if not more proficient. And at the Outlaw Open, the average max-rep set of pull-ups (after two miles of running and a 5-rep max deadlift) for the women was 33, just 9 behind the men.


Outlaw Open Notes:

The Outlaw Open is the first non-HQ sponsored competition that I've analyzed, so I thought it would be interesting to see how it compared to the Open, Regionals and Games, and also, what we can learn about the athletes who competed there (and there were some heavy hitters).

First, this won't come as much of a shock, but the Outlaw Open was heavy. Real heavy. The average relative weight on metcons for the men's competition was a 1.45*, which is 23% higher than the highest we've seen in the Games or Regionals. To give you a feel of what a 1.45 means, these are the types of weights you'd see in an average metcon:

Clean/Jerk - 195
Snatch - 145
Deadlift - 350
Thruster - 160
Overhead squat - 165
Back squat - 275
Front squat - 232 (we actually saw 250 in a metcon here)

The average weight for the women's competition was 0.95, which is higher than the average load in the men's Open in 2011. You may want to read that last sentence again to let it sink in.

In terms of LBEL, which takes into account max-effort events as well as the portion of the competition that was weighted/unweighted, the men's Outlaw Open came in at 0.94. That's just a shade below the 2012 Regionals, the highest in an HQ competition. The LBEL gives a better indication of whether a competition "favors" bigger athletes, so the inclusion of more bodyweight movements (in total the split was 50/50 between bodyweight and lifts) helped to balance things out. It's interesting to note that the men's CrossFit Invitational had an LBEL of 1.20, although that was a team competition so it's not exactly an apples-to-apples comparison.

As far as the types of movements involved*, here's how the Outlaw Open stacked up against the Open, Regionals and Games the past two years.


This kind of blew my mind that an Outlaw competition actually included Olympic lifts as a relatively modest portion of the scoring (I did take into account the points available on each event). So while the weight levels may have been most similar to the Regionals, the types of movements were distributed much more like the Games.

So what can we learn from the Outlaw Open? Well, clearly given the level of competition, anyone finishing well has to come away feeling good about their preparations for the upcoming season. I tend to think this will be most predictive of performance at the Regional level (vs. the Open or the Games), although we didn't see any lighter, higher-rep, grinding metcons, and the Regionals have had one of those the past two years. But if the Regionals are once again heavy, I think you'll see the top athletes from this competition with a good shot to get to the Games. Did someone say "vindication for Matt Hathcock?" We shall see...



*Assumptions on base weights for non-standard movements:
Wheelbarrow - 240 (same as deadlift)
Double KB thruster - 44 each arm (assumes DB's are 2.25x as difficult as a barbell, and KB's are 1.1x as difficult as a DB)
Double KB snatch - 40 each arm (same assumption)
Barbell step-up - 95 (this one was purely a judgement call based on watching the athletes a bit)
I'm always looking for more data to refine these base weights, so let me know if you have reason to believe these are out of line.

**If you've read my last post, you'll see that that I now have eliminated "Medicine Ball Lifts." I moved wall-balls into KB/DB lifts (I felt they are most similar to a DB thruster) and moved all other movements involving a medicine ball/slam ball (medicine ball cleans, ball slam, GHD ball toss, etc.) to "Uncommon Crossfit Movements."


Monday, June 25, 2012

Are certain events "better" than others?

Today's post will go in a slightly different direction than the previous three. Instead of focusing on the athletes, I'd like to focus on the events themselves and try to address the question (in the context of CrossFit competitions): "Are certain events better than others?"

Intuitively, I think most would agree the answer is yes. An event like, say, "Fran"will almost certainly test a person's fitness better than something like a competition to hit the longest drive with a golf ball. One way to look at this is using the purely CrossFit definition of fitness: of the 10 general physical skills (http://library.crossfit.com/free/pdf/CFJ_Trial_04_2012.pdf), I would say Fran hits on just about all of them (except maybe accuracy, balance and agility), while a long-drive competition hits on maybe three (accuracy, coordination and speed).

But this gets tricky to prove which events are better than other simply using the 10 general physical skills. Just in the above example, there is some wiggle room in saying which skills are tested even by those two events. So let me propose another definition of what makes a good test of fitness: a good CrossFit event will provide a strong indication of an athlete's ability to perform well in a wide variety of OTHER tests. To help explain, consider this example:


My contention here (and most CrossFitters would probably agree) is that "Elizabeth" is a better test of fitness than either a 5K run or a max bench press. For one, it tests more of the 10 physical skills, but I believe it does a better job of indicating which athletes would perform well in a wide variety of other tests. In this example, I assumed that there is generally no correlation between running a 5K and bench pressing. As such, I assigned random rankings to those events. However, a person who is strong on "Elizabeth" will probably do fairly well at both a 5K and a max bench press. So in this example, the person who has the best rank combined on both the 5K and bench press also has the top rank on Elizabeth.

This is an extreme example, and obviously I have rigged it, but it gives us an idea of how we can get a feel for which other events are good tests of fitness. The way we can do this is to look at the correlation between an athlete's finish on one event and their combined finish on a variety of other events. In the above example, the correlation between "Elizabeth" and the combined ranks on the other two is 98%. The correlation for the 5K run vs. the other two is 0%, and for the bench press vs. the other two it is -8%. What this says is that the bench and the 5K don't tell us much as much as "Elizabeth." All we need to test is "Elizabeth," because that tells us just as much as testing all three events. Note that we are talking purely about TESTING fitness here, not training for it. It might be the case that the 5K is worthwhile in training, but not as much in testing.

So on this theoretical basis, I decided to look at the events we have seen thus far in 2012. Like I did in my first analysis, I limited the field to athletes who completed all 6 events at regionals, which gives us a sample of about 250 men and 250 women. I used my adjusted regional results (see first two posts), and my measure of how well an athlete did on each event was simply the rank*. For each event, I looked at the correlation between an athlete's rank and his/her combined rank on the other 10 events.

Let's start with a visual representation. Here is a scatter plot of the men's Open WOD 3 ranks (x-axis) vs. the combined ranks on all other events.


It is fairly clear from this plot that a better rank on Open Event 3 (further left) corresponds to better results on the other events. Now let's look at the same scatter plot for men's Regional WOD 1 ("Diane").


Whoa. While there does appear to be some weak correlation, it's pretty clear that the results from "Diane" don't do much to predict how an athlete will do on the other events. I don't find this particularly surprising. For years, we have seen otherwise solid athletes struggle with handstand push-ups, and at the elite level, there is just no way to make up any ground on the deadlifts at a weight as light as 225. So basically we are testing handstand push-ups, which do tell us something about an athlete's overall fitness, but not much - certainly not as much as we can learn by testing 18 minutes of box jumps, medium load push press and toes-to-bar.

So which how do the events stack up in terms of correlation**? Well, here are the results, with women first and men second:



Well, would you look at that? Men's Open Event 3 had the highest correlation and Men's Regional Event 1 had the lowest. You'd almost think I chose those two graphs on purpose. It is clear, though, that for both men and women, Open Event 3, Regional Event 4 and Regional Event 2 were strong predictors of success across the board, while Regional Event 3 and Open Event 1 did not tell us as much. Regional Event 1 did have a somewhat higher correlation for women than for men, possibly because the event was not so blazing fast.

We can see another trend from this chart as well: events with more movements tend to be better predictors of overall fitness. While this is not surprising, I think it is an important point. Single-modality events simply do not tell us as much about an athlete as a couplet, triplet or chipper***. I do not believe we should eliminate them from competitions for this reason, but I do think that there should be some consideration to weighting these events less heavily. The Games struck a good balance last year, in my opinion, by grouping the single-modality events together into "Skills Tests," which didn't put as much weight on any one of those movements. I think giving a max effort snatch or an extremely heavy dumbbell snatch the same weight as something like Regional Event 4 may not be appropriate (somewhere, Chris Spealler is nodding his head right now).

This is certainly not a topic with one absolute right or wrong answer. I would be very interested to see other opinions, not only on what I have done, but also on what defines a "good" CrossFit event.

*Note: I also looked at this another way, which was to give each athlete a score on each event that was equal to the percentage of work done relative to the overall top score/time. For now, I will ignore those results because they are generally the same as these.
**To give some perspective to what these correlations mean, you can square the values to get the "r-squared." Men's Open Event 3, for instance, has an r-squared of 56%, while Men's Regional Event 1 has an r-squared of 20%. One rough interpretation of the r-squared is that it tells you how much of the variance in the other events' scores is explained by the event we are using as a predictor. So Men's Open Event 3 explains about 56% of the variance in the other events' scores.

***Yes, Regional Event 5 actually had two movements, but the double-unders didn't have much of an impact other than as a tiebreaker. You could also argue that Regional Event 3 was basically only one movement, too, since the impact of the running was negligible for most athletes.