Follow me on Twitter!


Monday, September 24, 2012

What to Expect From the 2013 Open and Beyond

Assuming we can expect the Open to begin in late February again in 2013, we are now officially closer to next year's Open than we are to last year's. It's time to stop looking back at the 2012 Games season and start looking ahead to next year. I've written extensively about the elite athletes, but most of us who follow the sport closely are competitors ourselves, and this post is designed to understand the qualification process from the standpoint of someone planning to compete this coming season. This isn't about predicting the winners, it's about knowing what to expect and where to place your focus.

Now, seeing as don't have a direct line to Dave Castro and Tony Budding, I don't know what they're thinking for next year. But we do have two years and 6 competitions of data that can help inform us about what they're likely to throw at us in a few months. Let's start by looking at the last two years from the simplest, and possibly the most useful, angle: what movements have we seen the past two years, and how often have we seen them. The following chart shows every movement tested in the past two years along with the weight given to each movement. As I've done before, for each workout, I break it down into the movements involved and give each "station" equal weight. For example, on Open WOD 3 last year, box jumps, toes-to-bar and jerk each received a weight of 0.33. On Open WOD 1, burpees received a weight of 1.00 since it was the only movement.

This chart is a great starting point for understanding what HQ is testing when they're testing for the fittest on Earth.


In case you were not aware, you better get your Olympic lifting in order if you want to be competitive in CrossFit. Including the jerk, the Olympic lifts were worth about 20% of all events in the past two years. That's not likely to change. We've seen snatch tested in each the Open, Regionals and Games both of the last two years.

What's also clear is that the pull-up is still important, as are an array of other bodyweight movements, including muscle-ups, burpees and toes-to-bar. Throw in running and double-unders, and we're up over 50% of the total weight (after counting the Olympic lifts). You've got to be good at everything, but those are the basics.

However, these include all competitions. Because of logistic restrictions and the relatively lower skill levels, the Open includes a much narrower list of movements. Here's what we have seen from the Open the past two years.


There may a couple of other movements thrown in this year, but not many. I highly doubt we'll see running or swimming, and some other staple movements like rowing, handstand push-ups and rope climbs are not likely for one reason or another (I'm still holding out hope for HSPU's, but I think the odds are slim). Even the swing hasn't shown up in the Open so far.

So the main takeaway here is that for the Open, you basically need to be able to Olympic lift and handle some basic bodyweight movements. You need to be able to do those things very well, but if your handstand walk isn't on point yet, you'll probably be OK.

But we can look deeper. This only tells us what movements we're likely to see. For the lifting movements, there is another aspect we need to consider: the loading.

To understand how "heavy" certain workouts are compared to others, we can't simply look at the loading in a vacuum. The same weight might make for a very heavy thruster but a very easy deadlift. To solve this problem, I started asking around my gym for max lifts on a variety of common lifts. I used the maxes of those at my gym, along with a bit of research into some of the elite athletes and my own personal experience, to develop relativities between the lifts. I'd love to get a bigger sample in the future and refine these numbers, but for now, we'll work with what we have.

Once I got these relativities, I was able to set a "base" weight for each lift. With these base weights, I could then compare the loads we saw in the Open (and the Regionals and Games) and get a feel for how heavy they really are. Along those same lines, we can see what types of weights we could expect, on average, this year. Based on the past two years*, here are the "expected" weights this year.


Keep in mind, those are just the averages. We've seen relatively heavier and lighter loads than those in the past. Based on the heaviest workout we've seen (the squat clean and jerk from 2011) to date, we can make an educated guess about the heaviest weights we might see this year. Adding on a 5% margin in case Castro gets crazy, here are the heaviest weights you can reasonably expect to see required in a workout (rounded to nearest 5).

Clean: 180 (men), 120 (women)
Jerk:  175, 115
Snatch: 135, 90
Deadlift: 320, 215
Thruster: 145, 95
Overhead squat: 155, 100
Back squat: 255, 170
Front squat: 215, 140

Now, I know we're mainly focused on the Open, but let's give these numbers a little perspective by looking at the Regionals and the Games as well. I've calculated the average relative weight seen at each competition for the past two years. Things got a bit tricky at the Games with movements like the sled push, so there were some judgement calls**. Nonetheless, this gives us an idea of the relationship between the various levels of competition. 

A 1.0 is equal to the "base" weights in the charts above (135-lb. clean, 100-lb. snatch, etc.). These charts include metcons only (I'll discuss the max effort lifts in a bit, but of course there have not been any in the Open).

UPDATE 12/10/12 - This graph originally had incorrect values for the Open (they were too low). The error was only in this graph, not in the underlying numbers mentioned elsewhere.



As you can see, the loading gets substantially heavier once you get beyond the regionals. This makes sense intuitively to anyone who's followed the Games season the past two years. The average weights you could expect to see in a men's metcon at the regionals include a 155-lb. clean, a 275-lb. deadlift and a 185-lb. front squat. Those are just the averages - we've seen much heavier.

Finally, let's combine the two concepts of loading and movements. What I was curious about understanding was just how much emphasis was placed on lifting, and lifting heavy, in each competition for the past two years. Consider the 2011 and 2012 regionals: based on the average weight load in metcons, the competitions were roughly equal in terms of load. However, as I thought about it, it seemed intuitive that the 2012 regional was much "heavier." After all, we saw 225-lb. cleans, 345-lb. deadlifts and 100-lb. DB snatches. Additionally, we had a max-effort snatch. The 2011 regional was heavy (315-lb. deadlifts, 135-lb. snatches, max thruster), but it didn't seem as heavy. 

The reason? The 2012 regional had more lifting, although not necessarily heavier lifting. And that's what really matters if we're trying to judge these things. When Chris Spealler, a little guy, says the regional programming was really tough for him this year, what he means is that there were a lot of lifts and they were awfully heavy.

The following chart shows the ratio of bodyweight movements to lifts in each of the last six competitions.


We can see that the 2012 Regional has had the biggest lifting bias, by a decent margin, of any competition in the past two years. Using this information and the average loading we developed earlier, we can get a total picture of how much heavy lifting was emphasized at each competition. The following metric, which I've called Load-Based Emphasis on Lifting (LBEL), is calculated by taking the average weight load (including max effort lifts***) and multiplying that by the percent of the movements that are lifts. The numbers here are harder to interpret, but this may help: "Fran" for men is about 0.45, "Grace" is 1.0 and "Cindy" is 0.0. Do all three of those for a competition, and you'd get about 0.5 for the competition as a whole.

Here are the LBEL scores for the past two years.


You can see that the 2012 Regionals was by far the top score. You may also notice that the Games are often equal to or lower than the Open. This is because the Games often focuses heavily on bodyweight movements not seen at earlier levels (long runs, obstacle course, deficit HSPU, handstand walks, etc.). The Games is not "easier" than the Open by any stretch. The loads you do see at the Games would be too much for 95% of the athletes who compete in the Open, and the strength demands of the bodyweight movements are extreme. But if you want to know which programming more favors a bigger, stronger athlete over a smaller, better conditioned athlete, I think they are actually pretty similar.

The main takeaways here are these: 1) For the Open, the emphasis is on Olympic lifting and some basic bodyweight movements; 2) The loadings at the Open are moderate, but lifting is emphasized quite a bit; 3) Things are likely to get heavy at the regionals; 4) The Games has more emphasis on bodyweight movements, but be prepared to lift heavy when you do lift; 5) It's time to get back to training - only 5 months left until the 2013 season begins!


Math notes:
*For the snatch workout this year, the weights varied based on how far you got in the workout. I looked at the top 1,500 overall finishers (roughly the regional qualifiers), and calculated an "average" load throughout the workout. The reason for using the top 1,500 is that I wanted this post to be from the perspective of someone attempting to qualify this year. The average weight came out to be 128 lbs. for men and 76 lbs. for women. In calculating the average, I looked at the total amount of weight lifted, then looked at how much of that weight was lifted at each level. For instance, if you got a score of 60, that's 30x75 = 2,250 lbs. lifted at 75 lbs. and 30x135 = 4,050 lbs. lifted at 135 lbs. Therefore, 64% was lifted at 135 and 36% was lifted at 75. The average would be .64x135 + .36x75 = 113.6 lbs.

**Here are the relative weights I assigned to the lifts not listed on the original "base" weight chart (these are the men's weights, women's were generally reduced to about 2/3 of these): 
Swings (53 lbs) .75, DB ground-to-overhead (45 lb DBs) 1.00, weighted lunge (45 lbs) 1.00, DB snatch (100 lbs) 1.50, farmer's walk (100 lb. DBs) 1.50, weighted pull-up 1.08, water jug carry 1.50, dog sled 2.00, sumo deadlift high pull (108 lbs) 1.14, sled pull with rope 2.00, ball toss (4 lbs) 0.50, medicine ball clean (150 lbs) 1.5, blocking sled push 1.50, sledgehammer 0.75. 
If you'd like more info on how I arrived at these, let me know and I'll expand on my thought process.

**For the max effort lifts, I took the average result for each competitor. Max effort lifts are tricky because they are dependent on who is doing the workout (similar to the Open snatch WOD). In general, I'd expect most max effort lifts to be about twice the average weight load seen at the competition. That would mean the metcons are generally done at about 50% of each person's max, which is pretty typical if loads are scaled properly.

Wednesday, September 5, 2012

Quick Hits: What Can We Learn from the Standard Deviation System?

After posting a lengthy essay on the benefits of switching to the Standard Deviation scoring system, I felt like I had to apply the scoring system to this year's Games and see what I could learn. Obviously it would be interesting to re-calculate the final standings with the new system, but I think there are a few other things that we can do with this new system. Because we are now scoring based on performance rather than simply rank, this new system allows us to compare events and individual performances across separate events. I'll try to keep this one relatively short, just hitting on the highlights of what I found.

Remember all events are converted to a power output (think reps or stations per minute, instead of time to completion). This was a painful process b/c of the way HQ scored athletes who did not finish a workout in the allotted time. I also generally assumed every "station" was worth equal weight (so on the Medball HSPU, the 8 medball cleans were equal to the 7 HSPU). This was the simplest solution. On the obstacle course, because I was basically forced into using the rankings and not the actual performances, I assumed a normal distribution and converted the ranks to an equivalent number of standard deviations from average. If HQ were to adopt this scoring system, they'd have to make the decision on how to weight the stations.

Anyway, enough with the math and onto the results:

Which event had the widest spread?

We can judge this based on the coefficient of variation for each event, which is the standard deviation divided by the average.

For the men, the widest spread came in the Medball-HSPU workout. The average score was 0.60 stations/minute (9:56) and the standard deviation was 0.18 stations per minute, giving us a coefficient of variation of 29%.

For the women, the widest spread also came in the Medball-HSPU workout. The average score was 0.56 stations/minute (which translates to finishing in 10:48, which is over the cap) and the standard deviation was 0.32 stations/minute, giving us a coefficient of variation of 58%. This shouldn't be surprising, considering the winning time was just over 5 minutes, but more than half the field didn't even finish.

Which event had the tightest spread?

For the men, the tightest spread came in the sprint. The average score was 6.60 meters/second (45.48 seconds) and the standard deviation was just 0.32 meters/second. The coefficient of variation was 5%. The second-tightest was Pendleton 2 at 8%.

For the women, the tightest spread also came in the sprint. The average score was 5.83 meters/second (51.49 seconds) and the standard deviation was just 0.31 meters/second. The coefficient of variation was 5%. The second-tightest was the clean ladder at 8%.

What was the most dominating individual performance over the field?

We'll measure this based on the number of standard deviations from the mean by the winner.

For the men, this came in the Rope-Sled event, where Matt Chan had a result of 1.32 stations/minute (7:33.6), which was 2.80 standard deviations above the average score of 0.87 stations/minute (11:48).

For the women, the most dominating performance came in the clean ladder, where Elisabeth Akinwale had a score of 235.6, which was 2.42 standard deviations above the average of 195.76.

What were the widest and tightest margins between first and second in an event?

Similarly, we're looking for the standard deviations between the first and second place finish.

For the men, the widest gap came in the Rope-Sled event. Chan was 0.89 standard deviations ahead of second-place Jason Khalipa, who had 1.18 stations/minute (8:27.0). The closest event came in the sprint, where Nate Schrader finished in 7.14 meters/second (42.0 seconds), just 0.11 standard deviations ahead of second-place David Levey at 7.11 meters/second (42.2 seconds).

For the women, the widest gap came in Elizabeth. Deborah Cordner Carson had a score of 25.09 reps/minute (3:35.2), which was 1.04 standard deviations ahead of second-place finisher was Kristan Clever at 21.79 reps/minute (4:07.8). The tightest race came in the clean ladder, where Lindsay Valenzuela basically tied Akinwale (she had the same lift but completed one fewer deadlift).

*In the events before any cuts, the biggest gap on the women's side came in the ball toss, where Cheryl Brost's score of 61 points was 0.53 standard deviations ahead of second-place Elizabeth Akinwale (57 points).

Seriously, what do the revised standings look like?

OK, I'm going to caveat this by saying that it's not totally fair to say the standings would have looked like this if we had scored the event differently. Obviously the athletes may have approached workouts differently had they been higher or lower in the standings, and they may have pushed harder for those extra few points in each event when more than just a simple ranking was involved. I think this is probably the least important thing we can learn from the new scoring system, since the event is done and there's nothing we can do to change it.

But that being said, for amusement purposes only, here is your revised top 12 for men and women (no one from outside the top 12 could move into the top 12 because of the cuts):

*UPDATED 1/19/2013 - An error in the Obstacle Course scoring has been fixed and these have been revised. Only major shift was Foucher going from 4th to 2nd on the women's side. Otherwise pretty similar to prior results. 
Men
1. Rich Froning (15.49)
2. Matt Chan (11.97)
3. Scott Panchik (8.75)
4. Jason Khalipa (7.99)
5. Kyle Kasperbauer (6.96)
6. Dan Bailey (6.57)
7. Austin Malleolo (5.27)
8. Marcus Hendren (4.82)
9. Nate Schrader (4.23)
10. Graham Holmberg (4.07)
11. Ben Smith (2.77)
12. Chad Mackay (1.78)


Women
1. Annie Thorisdottir (13.70)
2. Julie Foucher (9.19)
3. Talayna Fortunato (8.77)
4. Kristan Clever (8.55)
5. Camille Leblanc-Bazinet (6.14)
5. Lindsey Valenzuela (5.65)
6. Elisabeth Akinwale (5.60)
8. Valerie Voboril (5.28)
9. Jenny Davis (4.74)
10. Rebecca Voigt (2.84)
11. Stacie Tovar (1.43)
12. Christy Phillips (0.75)

*These scores assume that the ball toss, broad jump and sprint were given half the value of the other events.

Wednesday, August 22, 2012

Opening Pandora's Box: Do we need a new scoring system?

Until now, I haven't touched on what seems to be the most controversial topic when the Games roll around each year: the scoring system. I have been mainly focused on evaluating what happened this year and predicting results based on the system we have. But there is no doubt that the scoring system in place, which is based entirely on rank, has its flaws. The question is this: can we devise a system that is truly better?

Update: Before I get any further, I'd like to mention that the 2012 Open data I am using is from Jeff King, downloaded from http://media.jsza.com/CFOpen2012-wk5.zip. Thanks a ton to Jeff for gathering all the data. Much of this analysis would not be possible without it.

First, let me lay out four key flaws I see in the points-per-place system:
1) The results are heavily dependent on who you include in the field. Take the Games competitors, for example, and rank them based on their Open performance. You get very different results if you score each event based on the athletes' rank among the entire Open field than if you score each event based on the athletes' rank among only the Games competitors. Neal Maddox would move from 5th using the entire field to 2nd using only Games competitors, Rob Forte would move from 15th to 25th and Marcus Hendren would move from 34th to 16th. This is a problem.
2) There is no reward for truly outstanding performances. In the Open, Scott Panchik did 161 burpees in 7 minutes. The next closest mens Games competitor was Rich Froning at 141. In a field with only Games competitors, Panchik would only gain ONE point on Rich. He was not rewarded at all for any burpees beyond 142 (even with all Games competitors included, he only beat Froning by 33 spots, a relatively slim margin among 30,000+ competitors).
3) Along the same lines, tiny increments can be worth massive points if the field is bunched in one spot or another. If I had performed one more burpee (I did 104), for instance, I would have gained 857 spots worldwide. The difference between 70 and 71 burpees (a larger proportional increase in work output) was worth only 327 spots. And the gap between 141 and 161 was only 33 spots.
4) Other athletes can have a huge impact on the outcome between two competitors. Why should the differential between Rich Froning and Graham Holmberg come down to how many other competitors finished between their scores on a certain event? If those other competitors hadn't even been competing, it wouldn't change how Rich and Graham compared to each other.

Now, while I don't agree with everything Tony Budding says, I think he brought up a good point when he defended the scoring system on the CrossFit Games Update show earlier this year. Regardless of whether it has some mathematical imperfections, the fact of the matter is the points-per-rank scoring system is very easy to understand and very easy to implement. Watching the Olympic Decathlon, which has been refining its scoring system for years, reminded me why the points-per-place system isn't so bad. Unless you have a scoring table and a calculator handy, the Decathlon scores seem awfully mysterious. So if we're going to come up with a scoring system to replace the points-per-place system, I believe it has to be easy for viewers and athletes to understand.

That being said, we can learn from what the Decathlon has done. The idea behind the Decathlon scoring system is to attempt to weight each of the events equally, so that performances of equal skill level in each event yield similar point totals. Beyond that, the same scoring system should be applicable to athletes ranging from beginners to elite athlete. Additionally, the scoring system for all events is at least slightly "progressive" - this means that as performances get closer and closer to world record levels, each increment of performance more and more valuable. For instance, the difference in score between a 11-second to a 12-second 100 meters is wider than the difference between a 12-second and a 13-second 100 meters.

Each event is scored based on a formula, taking one of two forms

Points for running events = a * (b - Time)^c
Points for throws/jumps = a * (Distance - b)^c

For each, the value of b represents a beginner level result (for instance, 18.00 seconds in the men's 100 meters), and c is greater than 1, which is the reason the scores are progressive. Certain events are more progressive than others; generally, the running events are more progressive than the throws. Here is a chart showing the point value of times in the 100 meters.


The Decathlon scoring system, for all its complexity, generally does a good job distributing points among the 10 events. It also rewards exceptional performances much more so than the points-per-place system we use in CrossFit. However, there is simply no way to create such a system for CrossFit, even if we were fine with the complexity. Why? Because the events are unknown, and they almost always have never been performed before in competition, which means calibrating the formulas to be appropriate would have to be done on the fly. There was no objective measure about what a "good" performance was on the Track Triplet before it occurred this year's Games, and there certainly was no way to say what was an equivalent performance on the Medball Clean-HSPU workout, for example.

Of course, it's easy to pick apart other scoring methods, but the key question here is whether we can come up with anything better. In thinking about this post, I initially considered three types of systems: 1) a logarithm system, in which all performances are converted to logarithms, which gives us an indication of scores relative to one another; 2) a percentage of work system, where the top finisher is awarded a score of 100% and all others are scored based on their performance relative to that performance; and 3) a standard deviation system, where each finisher's score is based on how far from the average score they fell.

As we move away from a points-per-place system, there is one key point that need to be addressed. Since we are now considering differences in performance rather than just rank, we must think about how much a repetition of one movement is worth compared to a repetition of another movement. Think of Open WOD 4: one muscle-up is far more difficult than one double-under. If we count each movement equally, an athlete who completes ten muscle-up scores 250, which is only 4.2% higher than an athlete who completes all the double-unders but no muscle-ups (240). Clearly, this does not accurately reflect the difference in performance, and the movements need to be weighted accordingly. I think that it would not be too difficult for those designing the workout to make the points-per-rep system clear when the workout is announced. For example, HQ could simply say that each segment of the workout is weighted equally; completing 150 wall-balls is worth one point, completing 90 double-unders is worth one point and completing 30 muscle-ups is one point (10 muscle-ups is then worth 0.33 points). That's still a little light on the muscle-ups, in my opinion, but it is a simple solution for now, and it works well for most workouts (I'll use it for most events in my comparisons throughout this post). HQ could come up with whatever weightings they feel are appropriate. Sure, they would be somewhat arbitrary, but the workouts themselves are also arbitrary; if HQ lays out the rules, people will play by them.

Now, let me first discuss the logarithm system, which is definitely the most unusual of the three. The key point about logarithms, in this context, is that the difference in two athletes' scores is based only on the ratio of their performances. For example, let's say we had 3 athletes, one of which completed 40 burpees, one of which completed 80 burpees and one of which completed 160 burpees. The logarithm scoring system (we'll use a natural logarithm, although the base is irrelevant) would give athlete A a score of 3.689, athlete B a score of 4.382 and athlete C a score of 5.075. The difference between athletes A and B is .693, which is exactly the same as the difference between athletes B and C. By using this system, we can compare a 20-minute event exactly as we'd compare a 5-minute event: it's only the ratio between athletes that is important. The scores are also completely independent of who is in the field.

However, the logarithm system has a couple of significant drawbacks. First, it is certainly not easy to interpret, and most non-math majors might have a tough time recalling what a logarithm even is. But more importantly, this system does not reward the outstanding scores whatsoever. It actually does the reverse of the Decathlon's progressive system: as scores get better, you need a wider and wider gap in performance to gain the same point value. Scott Panchik's 161 burpees would give him the same point advantage over Rich Froning's 141 as an athlete doing 40 burpees would gain on an athlete doing 35.

So as mathematically pleasing as it is, let's drop the logarithm from the discussion. Let's move on to the percentage of work method. This method is simple: to score an athlete, we simply take the ratio of their score to the top score in competition (personally, I'd keep the genders separate). My score of 104 burpees would be translated to a score of 64.6% (104/161). Using this system, here are the top 5 men's Open results among competitors who reached the Games:


Note: In calculating the scores for each workout, I assumed each portion of workouts 3 and 4 were weighted equally (for WOD 3, 15 box jumps = 12 push press = 9 toes-to-bar = 1.00 points each). For workout 2, I weighted each rep by the weight used. The first 30 reps were worth 75 points each, then 135 each for the next 30, and so on. Workouts 1 and 5 were scored with all reps counting equally.

Keep in mind that for workouts with a set workload performed for time, we need to convert the times to work-per-unit of time. For instance, doing Fran in 4:00 could be converted to 90 reps/240 seconds = 0.375. A 5:00 Fran would be 0.300, which would be 80% of the work (well, technically power, not work) of the 4:00 Fran. If we have an event where all athletes might not finish within a time cap, we need to be careful to weight the reps appropriately (as described above). For instance, if Open WOD 4 had been prescribed as 150 wall-balls, 90 double-unders, 30 muscle-ups for time (12:00 time cap), we use our weights to accurately score all those athletes who did not finish in 12:00.

This method solves many of the issues we had with the points-per-place system. The only part of the scoring system that is dependent on the rest of the field is the winner, and most fields of competitors will have a winning score that is the same ballpark for a given workout. If you were to restrict the field to only Games competitors, the results would be identical. Outstanding performances are indeed rewarded, like Scott Panchik's 161 burpees (12% spread over next highest Games athlete). Bunching in one spot is not an issue, because athletes are scored based on performance only, not rank. Similarly, other competitors finishing between two athletes has no bearing on the relative scores of those two athletes.

However, there is one major concern about the percentage-of-work system. This method assumes that the athletes' scores will be distributed between 0 and the top score in a similar fashion for each workout, when in reality, some events are naturally going to have a tighter pack. Consider the sprint workout at the Games: the last-place competitor on the men's side would have received a score of 78%. On the medball clean-HSPU workout, the last-place competitor would have scored just 30%. Essentially, the sprint workout becomes much less meaningful than the medball clean-HSPU workout because there is much less opportunity for the winners to gain ground. This is easy to see when we compare the distributions of the two workouts graphically.




There are a couple of options to remedy this. One option is to modify the percentage-of-work system so that we see where an athlete's percentage of work falls between the lowest and highest score. Using this method, the 30% on the medball clean-HSPU workout and the 77% on the sprint both receive a score of 0%. A score of 65% on the medball clean-HSPU would score 50%, as would an 89% on the sprint workout. The problem with this solution is that one outlier performance can skew the low end. In the Open, using the entire field, there was a score of exactly 1 on every workout. Even among Games competitors, there may be one athlete who either is injured or simply struggles mightily with a particular movement, and that can drag the low end down unfairly.

The second option is to use the standard deviation system. This system looks at how far an athlete was from the average score in a given workout, taking into account how spread out the scores are. To calculate an athlete's score, we use the following formula:

Score = (Athlete's Result - Average Result) / Standard Deviation

For those unfamiliar with a standard deviation, it basically gives an indication of how far in either direction most athletes were from the average. If a distribution is normal (which most of these workouts tend to be), then in general, about 2/3 of the scores will fall within 1 standard deviation of the average. About 95% will fall within 2 standard deviations of the average. A related concept, called the coefficient of variation, tells us how large the standard deviation is compared to the mean (which basically indicates whether we had a tight pack or a more spread out field). The coefficient of variation for the sprint was 4.9%, but on the medball clean-HSPU event it was 28.8%.

On the sprint event, the average result was 6.59 meters/second (45.53 seconds). The winning speed was 7.14 meters/second (42.00) seconds. The standard deviation was 0.32, so the winning time would receive a score of (7.14 - 6.59) / 0.32 = 1.73. The worst time (5.51 meters/second, or 54.40 seconds) would receive a score of -3.36, giving us a total spread of 5.09. On the medball clean-HSPU event, the winning speed was 0.96 stations/minute (finished in 6:15.8). The standard deviation was 0.18, so the winning time would receive a score of 1.98. The worst time (0.29 stations/minute, or 10:00 plus 25 reps remaining) would receive a score of -1.83, giving us a total spread of 3.81, which is actually considerably less than the spread on the sprint workout. The reason is that the score of 54 seconds in the sprint was well outside the normal range, and it was punished accordingly.

Update: Using the standard deviation system (with all Games competitors included in calculating mean and standard deviation), here are the top 5 men's Open results among competitors who reached the Games:



Mathematically, the biggest drawback to this system is that it is somewhat dependent on the field. On Open WOD 1, the overall average score was 95.4 (among men under 55 years old) with a standard deviation of 17.2. If we limit that to only Games competitors, the average is 123.8 and the standard deviation is only 8.8. This makes an outlier performance like Panchik's 161 burpees more valuable when we only look at Games competitors than if we look at the whole field. Still, each competitor moved an average of just 1.3 spots in either direction when we switched the field from all Open competitors to Games competitors only. Using the points-per-place system, each competitor moved an average of 3.5 spots.

My feeling is that, despite this drawback, the standard deviation system is the optimal solution. I understand that the term "standard deviation" may sound foreign to many athletes and fans, but it is a relatively simple and intuitive mathematical concept. And we can easily change the name to something less intimidating, perhaps the "spread factor" or simply the "spread." If the weighting of each movement is clearly defined beforehand, the calculations for the scores of each workout should not be overly difficult, and the results should be fairly easy to understand. Certainly it would be far more transparent than the Decathlon system, while providing a similar level of fairness. There is also the convenient property that a total score of 0.00 is exactly average.

Imagine competing in the Open with this system. Once you have completed your workout, assuming there have been at least a few thousand entries so far, you already have a reasonably good idea of your score. The average and the standard deviation will not change much over the course of the next couple days. You won't need to worry about a logjam at one particular score unduly influencing your own result. The effects of attrition (fewer people completing the workout each week) should be basically negated, since we are not scoring based on points.

In my view, this is a much more equitable Open. It also makes for a more equitable Regional and Games competition. Does that mean HQ will veer from their hard-line stance on the points-per-place system? I have my doubts. But hopefully this provides some insight into why this is a discussion worth having.

Sunday, August 5, 2012

Were the Games Well-Programmed? (Part 2)

In this post, I'd like to look at the 2012 CrossFit Games season as a whole. In response to the question "Were the Games Well-Programmed?", it's going to be difficult for anyone to give an absolute "yes" or "no." Still, I think we can certainly look back and see aspects that were done well and other areas where I believe HQ could improve.

In my last post, I gave a generally positive review of the programming in the CrossFit Games finals. But the Games cannot simply be viewed alone, because the athletes competing were only there because of their performances in the Open and the Regional. To be sure, athletes could not have any glaring weaknesses, or else they would not have made the Games at all. But let's look at the programming across all three levels of competition and see where HQ put the most emphasis.

The following table shows every movement that was used in competition this season. As you can see, more than 30 distinct movements were tested, and very few, if any, CrossFit staples were left out. However, the extent to which the movements were tested varied widely. In adding up the total value assigned to each movement, I assumed that each workout was worth a total of 1.00 (Games workouts scored on a 50-point scale were worth only .50). Within each workout, I assumed that each "station" in the workout was worth equal value, so the box jumps in Open WOD 3 were each worth 0.33 points, whereas the burpees in Open WOD 1 were worth 1.00 points*.



What is clear from this is that HQ puts a large value on the Olympic lifts. The clean and snatch were worth a total of 5.35 events on their own! Add in shoulder-to-overhead (0.67) and that's more than 6 events worth of points based on the Olympic lifts. Although I am a big fan of the Olympic lifts myself, I do think the snatch in particular was over-valued. It was worth nearly 14% of all the available points, including 20% of the Open and 17% of the Regional. The pull-up, a CrossFit staple for years, accounted for 40% of the value of the snatch (maybe slightly more if you considered the pull-up-like elements of the obstacle course). 

However, in total, the lifting bias was not as great as some people believe. In total, purely bodyweight movements (excluding running, but including the obstacle course and double-unders) accounted for 45% of all available points; barbell or dumbbell-based movements accounted for about 38%; running or rowing accounted for 6%; all others (including medball lifts) accounted for 14%. I think there was good balance here, with the exception of the running and rowing. 

I think the lack of running in the Open and Regionals showed in the Games. For both men and women, neither of the run-focused events (shuttle sprint and Pendleton 2) were highly correlated with success across all other events in the season. In fact, the sprint had basically 0 correlation with success in all other events for the men. For comparison, two charts are below: one shows the weak correlation between men's shuttle sprint and all other events, and one shows the strong correlation between women's Open WOD 3 and all other events (the concept of correlation with other events is detailed in my post "Are certain events 'better' than others?").



In other words, the shuttle sprint was sort of a crapshoot, because the top finishers didn't necessarily do well in those events, whereas Open WOD 3 was dominated by athletes who did well across the board. My feeling is that because running was not tested earlier, we may have omitted some athletes who would have done better on the running events at the Games.

Let's look a bit more into the qualification structure on the road to the Games. The Open, Regionals and Games should all be testing similar things, and in my mind, there are two over-arching goals when programming and carrying out the Open and Region rounds: 1) In the Open, find the athletes with the best shot of reaching the Games, and 2) at the Regionals, find the athletes with the best shot of winning the Games. Put another way: 1) The Open should not eliminate any athletes who would have had a legitimate shot at reaching the Games if they had competed at Regionals, and 2) The Regionals should not eliminate any athletes who would have had a legitimate shot at winning the Games if they had qualified. It is certainly possible to disagree with that sentiment, but my feeling is that we want to pick the best athletes for the Games. We do not want to send athletes to the Games who will not do well there.

So, let's take a look to see if those goals were accomplished. It is impossible to say for sure how the eliminated athletes would have done, but there are ways to get a good sense. First, let's look at the lowest Open finishers to make the Games. On the men's side, Patrick Burke took 35th in his region (Southwest) and Brian Quinlan took 27th (Mid-Atlantic). For the women, Caroline Fryklund took 25th (Europe) and Shana Alverson took 22nd (South East). Given that no one below 35th (and hardly anyone below 20th) wound up reaching the Games, I highly doubt any athletes placing below 60 in the Open would have reached the Games. In this respect, I think the Open did its job. That being said, I think that with the size of the competition pool increasing so rapidly, expanding the Regionals beyond 60 (possibly 100?) might make sense, although logistically this might be challenging.

At the Regional level, it was well-documented on the Games site just how challenging it was for even the elite athletes to qualify for the Games. Notable former Games athletes like Blair Morrison (5th in 2011) and Zach Forrest (12th in 2011) were unable to qualify this season. Could these athletes, or others who narrowly missed out, have contended for the title? Again, it is impossible to know for sure, but we can use the cross-regional comparison to look at the odds.

Because of the points-per-place scoring system, the cross-regional comparison can vary slightly based on how large of a field we use, but I have used a scoring system that includes all athletes who completed all 6 events. I also adjusted for the week of competition (as detailed in my first two posts, a couple months back). Using this system, let's look at the highest finishers not to make the Games. On the men's side, we had Gerald Sasser (21st - Central East), Joseph Weigel (22nd - Central East), David Charbonneau (26th - North East), Nick Urankar (29th - Central East) and Ryan Fischer (30th - Southern California). On the women's side, we had Andrea Ager (19th - Southern California), Sarah Hopping (32nd - Northern California), Chyna Cho (33rd - Northern California) and Amanda Schwarz (38th - South Central).

Now, in the Games, let's see how well athletes with similar ranks in the regionals did. For men, the highest finisher to finish worse than 21st in regionals (i.e., worse than Sasser) was Chad Mackay, who took 9th at the Games despite ranking 32nd in this regional comparison. The next-highest was Patrick Burke, who was 16th at the Games and 24th in the regional comparison. So it is probably fair to assume that none of the non-qualifying athletes would have been able to challenge Froning for the title, but certainly they could have made a run at finishing in the top 10. For women, however, several top women finished lower than Ager in the regional comparison, including Jenny Davis (8th at Games, 28th at Regionals), Christy Phillips (11th at Games, 20th at Regionals), Deborah Cordner-Carson (13th at Games, 34th at Regionals) and Cheryl Brost (15th at Games, 21st at Regionals). Could Ager have challenged Annie Thorisdottir for the title? I doubt it, but given her Regional performance and her Open result (6th in the World), I think it is not out of the question that she could have challenged for a spot in the top 5.

I think the women's results do indicate that some top athletes might have missed the Games. Now, was this a result of poor programming at Regionals, or perhaps do we simply need more qualifying spots? In Ager's case, if we look at the athletes from her region who did make the Games, we see that all four (Kristan Clever, Rebecca Voight, Valerie Voboril and Lindsey Valenzuela) finished in the top 10, so this leads me to believe that the programming was not the issue. The bigger issue is that certain regions are simply too competitive. Consider the men's Central East: all five qualifying men finished in the top 10 (including the champion), and five other men were in the top 35 in this cross-regional comparison (the three mentioned above, plus Elijah Muhammad and Nick Fory). Other regions, such as the North West, had no athletes in the top 20 at the Games. I don't think it's unfair to suggest that HQ consider re-allocating the Games spots or adding more spots across the board.

Overall, I think we have to consider the 2012 Games season a successful one - the increased participation and interest in the Games speaks for itself. With that in mind, I believe there are clearly some adjustments that need to be made moving forward. Hopefully we see HQ continue to refine the system in 2013.



*Notes on valuation of movements: I broke down burpee-box jumps and burpee-muscle-ups into two movements, each worth half of that station's total value. For instance, in the Games Chipper, there were 11 total stations, one of which was burpeee-muscle-ups. So burpees and muscle-ups were each given 0.5/11 (~0.04) points. Also, I ignored the run portion of Regional WOD 3 (DB snatch/run) because it was virtually inconsequential to the results.

Tuesday, July 24, 2012

Were the Games Well-Programmed? (Part 1)

This is the first part of a two-part look at how well the 2012 CrossFit Games season was programmed. Today, I only want to focus on the CrossFit Games itself, ignoring the Regionals and Open for now. Along the same lines as my post "Are certain events 'better' than others?', this won't be a discussion with a clear-cut answer. What I'm hoping to do is take an objective look at the programming of the CrossFit Games, something beyond just "Whoa, Dave Castro is CURRRRAAZZZYY for making them do a triathlon!"


OK, the bulk of my post is based on the opinion that in programming the Games, there should be five goals in mind (in descending order of importance):


1) Ensure that the fittest athletes win the overall championship
2) Make the competition as fair as possible for all athletes involved
3) Test events across broad time and modal domains (i.e., stay in keeping with CrossFit's general definition of fitness)
4) Balance the time and modal domains so that no elements are weighted too heavily
5) Make the event enjoyable for the spectators


While this may not be HQ's stated mission in programming the Games, it's certainly how I would approach programming the Games and I think it is roughly in line with what HQ purports to do. With that as the framework for this analysis, let's evaluate how well Mr. Castro and HQ fared this year.


1) Ensure that the fittest athletes win the overall championship - Although I did pick Julie Foucher to edge Annie Thorisdottir for the women's title, I have to concede that Thorisdottir and Rich Froning again appeared to be the fittest athletes in the world, and by a fairly wide margin.


There is no question about Froning - he won the Open, he finished atop the Regional rankings (before and after adjustments based on week of competition) and he won the Games by 114 points. Unless Josh Bridges comes back next year or Froning gets hurt, Froning has to be a huge favorite to win it again next year.


As far as Annie is concerned, she finished third in the Open but finished atop the Regional (before and after adjustments based on week of competition) and wound up winning by 85 points at the Games. If we take all Games athletes and add up their rankings across the 21 events that everyone completed (excluding the five events after cuts started at the Games), Annie had 172 points, 25 ahead of second-place Julie Foucher. Keep in mind that she then beat Foucher in three of the final five workouts. (If you're curious, Froning finished with just 118, well ahead of second-place Dan Bailey with 226).


Adding in the fact that these two athletes both won the title last year, and I think it's safe to say that the fittest athletes did indeed win the titles this year.

Grade: A


2) Make the competition as fair as possible for all athletes involved - This goal involves a few different things. First, the scoring system needs to be fair. I think the scoring system is far from perfect, but I think it's fair enough. HQ is clearly trying to reward the particularly high finishes on each event by spacing out the top 6 places more, and I'm fine with that. I think there is an element of head-to-head competition that is rewarded by a system like this, as opposed to a pure place-per-event system like Regionals. Ideally, as many people have suggested, they'd switch to some sort of "percentage of work" system that would take into account the discrepancy between places (e.g. a 10-second victory is worth more than a 1-second victory on the same workout). But I think this system is OK for now.


Second, the events themselves need to be judged, operated and scored fairly. I think this was a mixed bag. The standards on finishing certain workouts, like the medball-HSPU and the Double Banger, were inconsistently enforced. Some athletes were made to get the Medball to stay in the required area after finishing the medball-HSPU workout, while others were allowed to simply drop the ball or merely run across the line. In some cases, this made a difference of 5-6 seconds. Also, the medball toss (without a doubt the worst event, more on that in Part 2) had several instances where the equipment malfunctioned (Kristan Clever is one example) and the athletes weren't given a fair chance. And although it likely wasn't a major factor, the athletes who went second in each heat on the medball toss had considerably less rest than those who went first (around 90 seconds vs. about 3 minutes). But as far as judging, I will say that from what I could tell, things looked pretty even, especially given that we're working mostly with volunteer judges.

All in all, I'd say this year's event was better than years past in this respect, but improvements still need to be made in this area.

Grade: B-

3) Test events across broad time and modal domains (i.e., stay in keeping with CrossFit's general definition of fitness) - To look at this, I've constructed the following table to compare all the Games events in terms of time, number of movements, the level of weight lifted and bodyweight strength required. The times here represent a ballpark figure of the fastest finisher for men and women. As far as level of weight lifted, I grouped everything into broad categories, generally based on the heaviest weight involved in the workout. Keep in mind, this varies by lift: for a Games athlete, a 200-lb. jerk is relatively heavy (at least in a metcon), but a 200-lb. deadlift is not. Bodyweight strength is based only on the non-weighted movements, such as handstand push-ups and toes-to-bar.




A couple of notes here: 1) The obstacle course was a hard one to pick as far as number of movements, so I went with three just to indicate that there were multiple skills involved. 2) The chipper might be considered medium as far as weight is concerned, but it seemed to have a definite strength bias if you look at the athletes who did well. 3) I considered the burpee-muscle-up to be two movements and the three separate sledge-hammer angles to be the same movement. These are debatable, to be sure. 4) I considered the clean to be 0 time, because although it technically took 5-6 minutes, 95% of that was rest. What they were measuring was the amount of work accomplished in that one-second clean, not how much could be done in 5 minutes.

Now, looking at the chart, we see times ranging from basically zero to more than two hours, but all but two workouts were under 10 minutes for the winner. I think they may could have done a bit more in the 15-25 minute range. As far as weight, I think they did a good job of mixing it up: by my count, 6 light-weight workouts, 5 medium-weight workouts, 4 heavy-weight workouts. They definitely tested bodyweight strength, as we saw three that included what I would consider "high" bodyweight strength movements (HSPU to a deficit, bar muscle-ups and regular muscle-ups) and two more that I would consider "medium" (rope climbs and ring dips). Most others included at least some bodyweight component. (I think you could argue that the rope climbs at 20' were comparable to muscle-ups, but I think the point remains.)

Although I did feel they were lacking in some of the moderately long workouts, going short enabled them to put athletes through 15 workouts and hit almost every common CrossFit movement. I think they did a good job overall in this area.

Grade: A-

4) Balance the time and modal domains so that no elements are weighted too heavily - This concept is pretty closely related to the prior one, but the focus here is on whether or not the different areas of fitness were fairly represented in the scoring. As previously noted, HQ definitely hit on just about every time and modal domain, but were certain areas over- or under-counted? We did count about that the number of heavy, medium and light workouts were fairly even, but the time domains do seem to be clustered a bit around the 0-5 minute range (9 workouts). HQ mitigated this a bit by assigning only 50 points to the broad jump, medball toss and sprint, but that doesn't entirely address the problem. 


Another way I decided to look at this was to see if the athletes' rankings for any of the events were particularly correlated with each other. If two events are highly correlated, then we may be giving extra credit to the areas of fitness that are tested there. Below are two charts showing all the correlations between the events for men first, then women (the final five events are excluded because not all athletes completed them):




There's a lot going on there, but I've highlighted some key cells. The yellow cells are combinations of events with a correlation coefficient above 0.60. The only times this happens is between Pendleton 1 and Pendleton 2 - this isn't surprising at all, because almost half of Pendleton 2 was Pendleton 1. There was definitely some double-counting going on here. However, I personally feel this is OK. The reason is that it is harder to test longer events as often because of the toll they take on the body, so giving the endurance guys two events that basically test the same thing makes up (to some extent) for the emphasis on short, explosive events later on.

The red cells show events that had a significant negative correlation. For the men, we see that event 10 (the clean ladder) was negatively correlated with Pendleton 1 - this is not surprising, considering endurance athletes generally lack high-end strength. For the women, those events were basically not correlated at all, but we did see a fairly strong negative correlation between Pendleton 1 and both event 3 (obstacle course) and event 4 (broad jump). Again, not that shocking because you're comparing an endurance event with two explosive, short duration events. I think it's fine to have events that are negatively correlated, and it's bound to happen given the nature of this sport.

For a quick comparison, let's look at this same chart for the men's decathlon at the U.S. Olympic Trials. There are only 16 athletes, but it gives us a taste at least. These correlations are based on the ranks of the athletes, although that is not the way the decathlon is scored:


You see there that there are four different instances of combinations of events that are highly correlated: 100m and long jump; pole vault and long jump; shot put and javelin throw; shot put and pole vault. The shot put/pole vault correlation is a little curious, but the other three make sense. For instance, top-end speed is a key factor in both the 100m and the long jump, so in effect, the decathlon ends up testing that type of speed twice. Keep in mind that this is only a sample of one decathlon meet, but it gives you an idea of how events can overlap.

So in all, I think HQ did a pretty good job in this respect. Again, I'd like to see them maybe scale back the short workouts and hit a few more in the 15-25 minute range, but overall I think they did well.

Grade: B+

5) Make the event enjoyable for the spectators - I'm curious to see how the coverage will be when condensed into 30-minute segments for cable, but I can say that overall, I was happy with the spectator experience in person. Other than the medball toss, in which the fans had absolutely no idea how well anyone did, and some of the early heats of the medball/HSPU, where hardly any athletes could complete the HSPU, I thought the events were enjoyable to watch. The set-up of the events in the arena was generally good, so that you could easily track the athletes' progress through the movements. The "hype men" did a solid job keeping everyone up to date on the athletes to watch, and they even got most of the pronunciations right! All in all, HQ did a good job setting this thing up for the viewer. I'll be interested to see what the response is from the general public when this thing airs in a few weeks.

Grade: A-


So that's it for today - next week, I'll try to tackle the entire season, including the Open and Regionals. Did we pick the right athletes to go to the Games? Which events did we like? Is this current set-up fair? We'll try to figure it all out next week. Thanks for reading!



Thursday, July 19, 2012

Initial Post-Games Thoughts

Before we get to the numbers, I have to say seeing the 2012 Games in person was truly a blast. I've watched online the past couple years, and we went to Regionals this year, but the experience at the Games was beyond my expectations. The professionalism of it all was very impressive, and although it's been said many times before, the crowd at the Games is unlike any other sporting event. Not louder or more energetic than any crowd I've seen, but by far the most friendly and congenial (not bad looking, either). If you have been on the fence about going in the past, definitely make it a priority next year. Grab your tickets early and get out to L.A.

Now, it's time to START to assess what went down last weekend. This will certainly not be the last post on the Games, and in fact, I still have several things I'd like to look into for the 2012 season as a whole. But for starters, let's see how our predictions panned out.

Notes: Just like I did in my 2011 analysis, I assigned points to each athlete who was cut for those events that they missed. The method of estimating those points is explained a couple of posts back.

Men: I honestly expected to do a little better here than I did. The model I used had an R-squared of 66% on the training data set (the 2011 Games), and while I obviously would not expect to do that well again, I expected to do a pretty decent job picking this year's Games. In the end, the model had an R-squared of 49% (that was calculated based on the actual points for each athlete, not just the rank). We did pick the winner correctly - Rich Froning won convincingly, as was expected. Matt Chan, on the other hand, had a performance that I simply did not see coming. His regional performance was solid (7th), but he was only so-so in the Open (21st), and at 33 years old, I didn't know whether he could handle the volume required these days in the Games. The model had him picked 23rd, and he ended up outperforming the model (in terms of points) more than any other athlete. As far as placement, Scott Panchik had the biggest move, finishing fourth despite being projected 27th by the model.

Despite the performances of newcomers Panchik and Marcus Hendren (7th), Games experience still proved to be a factor. Here is a comparison showing the average ranks of previous competitors vs. newcomers (same comparison we did for 2011 a couple of posts back):


Except at the top end, prior Games competitors outperformed newcomers with similar regional rankings. Something we may want to consider is modeling the top athletes slightly differently than the ones finishing more modestly at regionals.

What was really NOT much of a predictor this year was Open results. After taking into account the regional results, the Open results basically told us nothing more (in fact it had a slightly negative coefficient in the regression). The results did show that age was still a factor (younger athletes are expected to do better), but not to the extent we expected. Each year over 26 was worth somewhere around 7 points, after accounting for prior experience, Regional and Open results - the model had assumed more like 20 (after scaling up to the 1,350 available points this year from 1,000 last year).

Overall, it looks like we would have been better off simply using the Regional rank. The R-squared there would have been was about 55% using my adjusted regional rankings and 53% using the raw regional rankings. With another year of data under our belt, hopefully we can do better next year.

Women: The women's model was much simpler than the men's, and it actually slightly outperformed the men's as well. The R-squared (using points) was 50%, and of the top 10 women, I had 6 predicted to be in the top 10. If we had simply used the adjusted regional ranks and converted them to Games-style points, then used that to predict the Games points, the R-squared would have been only 46%. So taking the Open results into account certainly helped out.

At the top, as expected, it was a dual between Julie Foucher and Annie Thorisdottir. While we expected some other top names, like Kristan Clever (predicted 3rd, finished 4th) and Camille Leblanc-Bazinet (predicted 4th, finished 6th), Talayna Fortunato was a surprise. She was 11th in the adjusted regional rankings and 8th in the Open, but a 3rd place finish was unexpected (predicted 10th). Overall, Jenny Davis outperformed the model more than anyone else, finishing 8th despite being picked 26th.

For the women, we did see a bit more correlation between prior Games experience and improved results this year (last year we saw basically none). Here is the same chart as above, except for women:


What stuck out, however, is that the Open results were a pretty darn good predictor of success for the women. After accounting for Regional results, the regression actually showed that the Open was a slightly stronger predictor than the Regionals. If you had used the Open results alone to predict the Games, the R-squared would have been 47%, slightly higher than using the Regional results alone. And again, like 2011, age did not appear to be a factor at all once you take into account regional peformance (see 40+ year-olds Becky Conzelman and Cheryl Brost finishing 14th and 15th, respectively). Obviously, in general, it helps to be younger, but if a woman has already qualified for the Games, there is no reason to believe their age will negatively impact them any more in the Games than it did at Regionals or in the Open.


I'll continue to dig into this in the coming weeks. Hopefully everyone enjoyed the Games. Only 7 months until the 2013 Open begins!

Monday, July 9, 2012

2012 CrossFit Games - Who Ya Got?

OK, it's time to get down to business and predict some friggin' results, people. For background on how we got here, see the previous post ("The Method to the Madness that is Predicting the CrossFit Games").

First, for the ladies. The model I used was as follows: Games Points = 0.82 * Regional Points + 0.48 * Open Points + 123. Remember, points are calculated using the current Games scoring system (all events using the 100-point scale). For those of you saying, "That's boring - why did you only include those two variables and not, say, age or prior Games experience?" read the previous post. Now, if you're mentally prepared for the results, the table is below.


That, my friends, is pretty much a straight-up "pick 'em" at the top. Would I question anyone whatsoever for picking Annie (or for that matter, any of the top 6 women)? Absolutely not. BUT I AM NOT PETER KING, AND I STICK BY MY PREDICTIONS: Julie Foucher will win the CrossFit Games. Don't let me down, Julie.

Now, for the men. Things were a bit more complex (and hopefully more accurate for that reason). The model here is: Games Points = 0.56 * Regional Points + 0.41 * Open Points + 162 * Prior Experience - 15.68 * Age Beyond 26 + 174. Prior Games experience is a 1 for "Yes" and 0 for "No." So for the men, you should do well if you are young, have competed at the Games before and did well at Regionals and the Open. Drum roll please...


Folks, there is just no way to take the data we have from this year and NOT predict Rich Froning to win the CrossFit Games. The man won the Open and had the top Regional performance overall, he's only 24 AND HE WON THE GAMES LAST YEAR. Do other guys have a shot? Certainly - see my post "So who CAN win the CrossFit Games?" But Froning is absolutely the man to beat. 

You'll notice that the top projection for a Games rookie is Kenneth Leverich at 22. I would guess that we'll see a rookie finish higher than that, but I think it's unlikely we'll have another Josh Bridges come in and finish on the podium on his first try. 

So there you have it. Are these predictions going to materialize perfectly? Of course not. Are some of the predictions debatable? Yes (I mean Patrick Burke at 33 seems pretty low even to me). The R-squared for the women's model, using last year's results, was 56%, so that's basically saying 44% of the variance is not explained by this model. For men, it's about 66%. That tells you how difficult it is to predict the Games.

But on the whole, given the data we have, this is what I'm going with. Think you can do better? By all means, post your top 10 or even your top 50 and we'll see how it all shakes out.

ENJOY THE GAMES, EVERYONE!