Saturday, September 22, 2012

The OPSBI Fallacy

Yes, I realize I'm a year too late to this one.  But the MVP talk is a-startin', and I just know some idiot sportswriter is going to reference this "stat" when explaining their vote, so I thought it was worth another look.  I realize it's been torn apart by saberists before.  But I wanted to take a different look at OPSBI.

For those who don't know, OPSBI was created by Jim Bowden, the former Reds/Nats GM.  The basic idea is this:  you take the players on-base percentage, and add his slugging percentage (thus, the OPS).  Now, eliminate the decimal point (or multiply by 1000, however you prefer to think about it).  Then, add the player's RBI.  That's it.

So, first, I want to point out why this is stupid.  First of all, it doesn't measure anything.  At all.  It arbitrarily adds a rate stat (two rate stats, actually) to a counting stat, without any consideration for why.  It doesn't correlate to anything.  As we know, OPS has a reasonable correlation with run-scoring.  RBI represent actual runs batted in (though not runs scored).  So, you're adding something which correlates to run-scoring with something that actually is one-half of run-scoring.  Now, someone clever could probably make the argument that if you added Runs to this, you'd solve some of that problem.  But here's the question - why add anything to OPS at all?  I don't get it.

Anyway, Bill James says, "For a statistic to have value, it has to be meaningful with reference to something other than its own formula" (The New Bill James Historical Abstract, in the player comment on Craig Biggio).  OPSBI fails that test.

Of course, there are still defenders out there.  I don't feel like looking for articles right now, but I remember reading at least two of them last offseason.  Here's the thing they'll say, more or less:  "Who cares if it doesn't measure anything - it gives the right answer!"  Okay, well, in my mind, it gives the "right" answer - that is, it affirms (much of the time) the conclusion sportswriters have already made.  But I personally believe that I can throw out a lot of other BS stats that will do the same thing (more or less).  Anyway, I'll be looking at the last five years of data (not including 2012, of course, since we're still underway), and looking at the top MVP-finishers (non-pitchers only) for each season, and comparing them by different metrics.  Those metrics are:

MVP Finish - where did the player finish in subjective MVP voting?
OPSBI - of course.
RunAvg - (R+RBI)/AB
ButTheKitchenSink*Games - Games*(RBI+R+TB+BB+HBP+SB)/(3*AB) ; I already posted about this before.
 ButTheKitchenSink(rate) - (RBI+R+TB+BB+HBP+SB)/(3*AB) ; same thing, but as a rate stat (scaled to batting average); really, what this is, is RunAvg+BattingAvg+SecondaryAvg, and divided by three so that it looks like the players batting average.
StolenHomes - SB+HR
TripleCrowns - (3*HR)+((1000*Avg-100)/2)+(RBI)
HitByBallSacs - because I couldn't resist:  HBP+BB+SF+SH
rWAR - because, wouldn't it be fun if I used a metric that actually seemed to represent real value?

One last note before presenting:  in 2008, Manny Ramirez only played 53 games in the NL, so only his NL stats are included.  Also, I hope this formats okay.  Here come the charts:

2011 NL MVP OPSBI RunAvg BTKS BTKSr StoHos 3Crowns BallSacs rWAR
Braun 1 1105 .391 57.9 .386 66 343.0 66 7.7
Kemp 2 1112 .400 63.7 .395 79 366.0 87 7.8
Fielder 3 1101 .378 62.2 .384 39 345.5 123 4.3
Upton 4 986 .326 54.2 .341 52 294.5 82 5.7
Pujols 5 1005 .352 50.0 .340 46 322.5 72 5.1










2011 AL MVP OPSBI RunAvg BTKS BTKSr StoHos 3Crowns BallSacs rWAR
Ellsbury 1 1033 .339 54.9 .347 71 329.5 69 8.0
Bautista 2 1159 .405 64.6 .433 52 340.0 142 7.7
Granderson 3 1035 .437 62.3 .400 66 332.0 108 5.3
Cabrera 4 1138 .378 62.3 .387 32 337.0 116 7.3
Cano 5 1000 .356 52.1 .327 36 325.0 58 5.2










2010 NL MVP OPSBI RunAvg BTKS BTKSr StoHos 3Crowns BallSacs rWAR
Votto 1 1137 .400 60.4 .403 53 349.0 101 6.7
Pujols 2 1129 .397 63.6 .400 56 358.0 113 7.3
C. Gonzalez 3 1091 .388 53.3 .367 60 353.0 49 5.8
A. Gonzalez 4 1005 .318 52.8 .330 31 362.0 101 4.1
Tulowitzki 5 1044 .391 44.6 .365 38 306.5 59 6.5










2010 AL MVP OPSBI RunAvg BTKS BTKSr StoHos 3Crowns BallSacs rWAR
Hamilton 1 1144 .376 49.6 .373 40 343.5 53 8.4
Cabrera 2 1168 .432 61.4 .409 41 366.0 100 6.1
Cano 3 1023 .339 52.3 .327 32 326.5 70 7.8
Bautista 4 1119 .409 66.3 .412 63 362.0 114 6.6
Konerko 5 1088 .365 54.1 .363 39 345.0 83 4.3










2009 NL MVP OPSBI RunAvg BTKS BTKSr StoHos 3Crowns BallSacs rWAR
Pujols 1 1236 .456 72.6 .454 63 392.5 132 9.4
H. Ramirez 2 1060 .359 53.9 .357 51 325.0 76 7.1
Howard 3 1072 .399 59.5 .372 53 370.5 87 3.5
Fielder 4 1155 .413 65.9 .407 48 382.5 128 6.0
Tulowitzki 5 1022 .355 54.6 .362 52 304.5 85 6.3










2009 AL MVP OPSBI RunAvg BTKS BTKSr StoHos 3Crowns BallSacs rWAR
Mauer 1 1107 .363 50.9 .369 32 3028.5 83 7.6
Teixeira 2 1070 .369 56.7 .363 41 346.0 98 5.1
Jeter 3 937 .273 46.3 .302 48 269.0 82 6.4
Cabrera 4 1045 .326 53.4 .334 40 333.0 74 4.7
Morales 5 1032 .343 50.8 .334 37 329.0 56 4.0










2008 NL MVP OPSBI RunAvg BTKS BTKSr StoHos 3Crowns BallSacs rWAR
Pujols 1 1230 .412 63.5 .429 44 368.5 117 9.0
Howard 2 1027 .411 59.0 .364 49 367.5 90 1.5
Braun 3 994 .324 49.3 .326 51 322.5 52 4.3
M. Ramirez 4 1285 .476 25.3 .478 19 285.0 42 3.4
Berkman 5 1092 .397 62.9 .396 47 320.0 111 6.6










2008 AL MVP OPSBI RunAvg BTKS BTKSr StoHos 3Crowns BallSacs rWAR
Pedroia 1 952 .308 48.1 .306 37 280.0 73 6.8
Morneau 2 1002 .363 53.7 .330 23 325.0 89 3.9
Youkilis 3 1073 .383 52.9 .365 32 329.0 83 6.0
Mauer 4 949 .341 46.4 .318 10 267.0 97 5.3
Quentin 5 1065 .408 50.8 .391 43 316.0 89 5.1










2007 NL MVP OPSBI RunAvg BTKS BTKSr StoHos 3Crowns BallSacs rWAR
Rollins 1 969 .325 53.5 .331 71 314.0 62 6.0
Holliday 2 1149 .404 60.2 .381 47 379.0 77 5.8
Fielder 3 1132 .398 63.2 .400 52 363.0 108 3.4
Wright 4 1070 .364 60.4 .377 64 329.5 107 8.1
Howard 5 1112 .435 59.2 .411 48 364.0 119 2.8










2007 AL MVP OPSBI RunAvg BTKS BTKSr StoHos 3Crowns BallSacs rWAR
Rodriguez 1 1223 .513 73.6 .466 78 421.0 125 9.2
Ordonez 2 1168 .430 60.9 .388 32 376.5 83 6.9
Guerrero 3 1075 .373 53.1 .354 29 341.0 86 4.3
Ortiz 4 1183 .424 62.6 .420 38 353.0 118 6.1
Lowell 5 999 .338 48.2 .313 24 324.0 64 4.6

So, what often happens with a chart like this is, people either skip it and wait for the conclusion, or they read it and go, "so what?"  If you're in the former group, that's annoying, because charts take the most time to make.  So authors are upset at you for not reading the chart.  If that's you, go back and look at them.  We'll wait.

Okay, now that everyone's caught up, what the heck does any of this mean?

Well, if the goal of OPSBI is to correlate to a player's true value, we have to ask, "how do we best measure a player's true value?"  Thus, the first and last columns of the chart.  The last, WAR, is a statistical measure.  The first, MVP voting, is a completely subjective measure.  If OPSBI is so good at perceiving value, we'd really like it to have some correlation to one of these two or the other.

The problem is, it doesn't.  Not at all.  So either both the MVP voters and WAR are wrong, while OPSBI is the true best measure... or OPSBI is crap.  Now, I'll admit freely that both WAR and voters take defense into account, which OPSBI doesn't... still though, that's not (for the most part) how MVPs are won, or how WAR is decided, since offense bears so much more weight.

For instance, if OPSBI is such a good measure, why would Jacoby Ellsbury have finished above Jose Bautista last year in MVP voting?  Bautista was the better offensive player... even by these made up metrics.  I just don't get what OPSBI is doing that different from... ANY of these other things I made up.  Seriously.  Well, maybe not HitByBallSacs, but that's just too hilarious to not include.  Anyway, the others do just as good of a job predicting WAR and predicting MVPs... maybe better.  So why OPSBI?  I see no reason, since it doesn't even reflect voter tendencies.

Monday, July 23, 2012

Penn State

My fellow blogger here, Jordan, is a big PSU fan.  So why am I writing this?  Because if you hear stuff from a PSU fan, it's easy to ignore.  So here it goes.

I'm putting this paragraph at both the top and bottom of this article.  DO NOT accuse me of supporting Joe Paterno, or minimize his complicity in this horrible, horrible tragedy.  There are victims, who are/were CHILDREN, and they deserve all the protection in the world.  Ultimately, the rest of this nonsense, in which people yell at each other about statues and legacies and everything else is nonsense, and a waste of time, because what matters is that children were harmed in the most heinous possible way.  It's disgusting, Joe Paterno was somewhat complicit, and Jerry Sandusky is a monster.

Now, after writing how talk about anything other than the children is nonsense, I'm going to talk about something other than the children.  Why?  Because I'm not talking about Joe Paterno's "legacy."  I'm, in fact, not talking about Joe Paterno at all.  You know why?  Because I'm talking about the penalties levied against Penn State.  Because I have a question:  how does any one of these penalties help?

Okay, first, let's talk about colleges.  Colleges and Universities do a really good job protecting rapists.  That's true.  Ask anybody who's ever been sexually harassed/assaulted on a college campus.  Their attacker(s) still walk around campus with a slap on the wrist.  Schools usually protect students from legal repercussions if they drink underage or to excess, if they vandalize a building, if they violate a curfew, if they traffic in illegal substances, etc., etc., etc.  For those who make a stupid decision to smoke some weed, get caught, and avoid serious legal recourse, this is a wonderful thing, as those people often turn out to be productive members of society, glad they don't have a silly little black mark on their record for no reason.  But for sexual criminals, it's a bit more serious.  So there's a problem with the system in colleges and universities as it relates to sexual crime.

Second, let's talk about our system of criminal justice in this country.  Part of what it's supposed to do is rehabilitate.  Now, whether or not it does that is neither here nor there.  But it performs another function (theoretically, mind you).  It separates criminals from their would-be victims.

So let's look at PSU.  Well, the dangerous people (Sandusky, Paterno, etc., etc.) are gone from the university.  The second function of justice is served.  But what does taking away football wins, removing scholarships and Bowl games, and the rest of it (money aside, for a moment) do?

Nothing.

In fact, it trivializes the victims.  It says, "Here.  This ought to make you feel better.  Remember that 2003 Penn State team that went 3-9?  Well guess what?  Now, they went 0-12.  Doesn't that make the nightmares and disgusting feelings and difficulty having a normal sexual relationship for the rest of your life feel better?  Because it should.  Oh, and now some kids who have a chance to get a college education won't be able to, because PSU can't give them scholarships.  So that should help."

Uh... anyone else see the problem with that logic?  It's demeaning.  Now, taking down a statue so that no kid has to be reminded of a man who was complicit in (potentially) their own sexual assault as a child makes sense to me.  But the rest of it?  It's vengeance for vengeance's sake.  It's pure lynch-mobbing.  Societally, people feel better because PSU is punished.  But the fact of the matter is, there could have been more (and better) things done.

What would be better, you ask?  Well, give a bunch of money to the victims, first of all.  Obviously, that's trivializing, too, but there's no real restitution you can give, and at least that's something.  The second thing you do is PSU seriously converts major parts of its social sciences departments to studying victims, victimology, and helping develop new and helpful therapies for people.  Then you publish the work, and you help get it out there.  Then, PSU works with the NCAA to find out what the holes were in the Sandusky situation, and how they can be fixed when it comes to vetting college coaches.  How often are the examined by a psychologist/psychiatrist?  They'll be working with young adults.  Sure, they're adults.  But, as someone who's going to school to be a pastor, we learn that because there's a gross power imbalance (perceived or real) in our relationships with our parishoners, we have to be extra cautious, because they become a vulnerable population where our actions are concerned.  Likewise coaches.  Next, PSU works to find the other holes in reporting.  How do we make sure that a report doesn't get lost in a pile of papers and bureaucracy?  Automatic reminders to the complaintant?  A dual-reporting system, so that two organizations follow up on everything?  Whatever it is, ask PSU (and perhaps other Big Ten schools, or other Pennsylvania schools, or whoever) to research and implement a plan (or some plans) and test its/their efficacy.  Finally, PSU works with its criminal justice department and education department on profiling and/or reporting.  Workshops and curricula are developed to help kids talk about good-touch/bad-touch; they learn who to contact; they learn how to say "no," and that, when it comes to touch, no one is allowed to touch them (I don't mean this to sound like I'm blaming the kids for getting touched - what I'm saying is that kids are often really, really scared to report touching, and even if they can't say "no" in a moment, if they can tell a grown-up later, it can be a tremendous help).  And if they already have programs that teach them that, we look into how we can do it better.

Those are positive, active changes that a University holding some of the nation's best and brightest minds should be able to do.  Don't take away football wins, not because it's punishing for "sins of the fathers," but because it's insulting to victims.  Do the right thing, and work to change things for the better.  Make the world a better, safer place for kids everywhere.  They're the victims, and they're the ones we should be focused on.

I'm putting this paragraph at both the top and bottom of this article.  DO NOT accuse me of supporting Joe Paterno, or minimize his complicity in this horrible, horrible tragedy.  There are victims, who are/were CHILDREN, and they deserve all the protection in the world.  Ultimately, the rest of this nonsense, in which people yell at each other about statues and legacies and everything else is nonsense, and a waste of time, because what matters is that children were harmed in the most heinous possible way.  It's disgusting, Joe Paterno was somewhat complicit, and Jerry Sandusky is a monster.

Friday, February 17, 2012

Gary Carter and Catchers

Hello again, everyone.

There are a whole lot of tributes out there to Gary Carter today, and justifiably so. These tributes range from the purely statistical to the sappy and emotional; career highlights to personal remembrances; there are those who want to take this time to think only on the happy memories, and those who want to put the full man together in order to gain perspective. Grief is a complicated thing, and it makes sense that people would have so many different responses to Carter's passing.

It seems that much of the main-line baseball blogosphere is filled with people in their 40s and 50s (with some in their 30s), which makes sense, because they have had time to establish themselves. These are also the people on whom Gary Carter had the biggest impact, as they saw him as a young phenom in Montreal, and they saw him win a World Championship as a Met - and he was one of the major players in the famous game-6 rally.

But the 1986 World Series ended 11 days before I was born. Rob Neyer (click on "Gary" above to read the piece) wrote this morning that, in 1999, only 33% of the BBWAA voted for Carter for the Hall of Fame. I remember that same year worrying that this guy who was gaining momentum, Gary Carter, may prevent Robin Yount from entering the Hall on the first ballot, and that would just be stupid, because, you know - he was Robin Yount. He had to be a first ballot guy, because he was one of the greatest players of all-time (full disclosure: I do agree with that statement, but growing up in Milwaukee may give one an overly optimistic view of Yount). I didn't want this Carter - whoever he was - taking that away from Robin.

But, of course, hindsight is 20/20.

As it turns out, it was Yount who took votes away from Carter, and Robin was elected while Gary had to wait. Both were great players, and both were deserving of first-ballot induction, if you ask me. And that's where we come to. Because my greatest memory of Carter is somewhat antagonistic, I think I'll stick to the stats. Last month, I revealed my new statistical measure for Hall of Fame worthiness: WARSCOR. If you want more details, read the post. Needless to say, perhaps, what I want to discuss is Gary Carter, and his WARSCOR. Here are the top 9 catchers of all-time, and their WARSCORs (if you're wondering why nine, it's because those are the legit full-time catchers who have an argument for being the best ever, in my opinion):

Bench - 56.7
Carter - 52.5
Berra - 52.3
Piazza - 48.3
Dickey - 48.1
Cochrane - 46.2
Fisk - 44.8
Hartnett - 43.8
Rodriguez - 43.4

Keeping in mind that I only used gWAR (or The Baseball Gauge's WAR system), I thought that maybe I should try the other WAR systems out there. So that's what I did. Here they are, by rWAR, fWAR, and WARP, respectively:

Bench
- 52.9
Carter - 50.5
Piazza - 48.0
Rodriguez - 46.0
Fisk - 45.0
Berra - 44.6
Cochrane - 40.5
Dickey - 39.8
Hartnett - 34.5

Bench - 60.1
Carter - 53.6
Piazza - 52.4
Berra - 50.3
Rodriguez - 48.7
Fisk - 48.6
Dickey - 46.1
Cochrane - 43.6
Hartnett - 38.1

Bench - 56.5
Piazza - 52.9
Carter - 49.0
Berra* - 46.8
Fisk - 46.0
Rodriguez - 38.3
Dickey - N/A
Cochrane - N/A
Hartnett - N/A

*WARP data is only from 1950 on, so does not include Dickey, Cochrane, nor Hartnett; nor does it include Yogi Berra's first four seasons (1946-1949).

So, if you buy into my method (which, let's face it, you should, as I think it's an improvement on JAWS, CAWS, or wWAR), by 3 of the 4 WAR systems, Carter ranks as the #2 backstop of all-time, the only exception being WARP, which has him at #3. For reference, JAWS has him at #4, wWAR has him at #2, and CAWS has him at #4. The fact that Carter was on the ballot before Fisk, but got in after is pretty inexplicable in light of this information - especially when all of these systems rank him above Fisk. I'm not trying to denigrate Fisk - just point out that it's very, very reasonable to conjecture that Gary Carter is the 2nd greatest catcher ever. Rest in Peace, Gary. Your fans will miss you. And baseball will miss you - the second greatest catcher of all-time!

Monday, January 30, 2012

What Is a Compiler?

As you'll note from my last post, I've been playing around with a group of 500 players and assessing their worthiness for the Hall of Fame. One of the things that has always irked me about the Hall is when people talk about "compilers." Seriously - what's a compiler? Someone who sticks around forever and keeps putting up decent stats. Why is that bad? Isn't it good to have a long career? Well, of course it is. I do understand the perspective that if a guy gets 100 hits for 30 years to get to 3000, it's not the same as a guy who got 200 hits for ten, then 150 for 5 years, 100 more, then retired and fell short of 3000. But usually, when people argue against "compilers," it's really just a coded argument for, "The stats don't match my opinions, so I'll throw out a derogatory term to slight the player, instead of reconsidering my underlying assumptions." And that's really not what you want, especially from Hall of Fame voters.

Anyway, I don't just want to rant against Hall voters (though, really - who doesn't enjoy that?). What I'd like to propose is a quick mathematical model to see who the "compilers" really are. As in my last post, I'm using WAR as calculated on The Baseball Gauge. The method is simple: Take "peak" to be the player's ten best seasons. Then do some quick division: peak/career. Since I already had a pool of Hall-of-Fame-type players, I thought to just do the calculation quickly on them. I looked at the bottom thirty players - those for whom peak value was 68.4% or less of their career value. Here they are, in descending order:

Warren Spahn
Al Kaline
Lou Whitaker
Gaylord Perry
Stan Musial
Red Ruffing
Honus Wagner
Phil Niekro
Pete Rose
Bert Blyleven
Jack Quinn
Mel Ott
Greg Maddux
Rickey Henderson
Tris Speaker
Frank Robinson
Hoyt Wilhelm
Jim O'Rourke
Barry Bonds
Tommy John
Babe Ruth
Willie Mays
Ty Cobb
Roger Clemens
Dennis Eckersley
Don Sutton
Nolan Ryan
Cap Anson
Cy Young
Hank Aaron

That's right. The #1 compiler of all time is . . . Hank Aaron? Well, actually, it makes a lot of sense. Aaron was lauded for his consistency as a player, even when you adjust for ballpark, era, etc. So only 59% of his career value is wrapped up in his peak. Likewise, the next three players had such long and excellent careers that they can't be blamed for putting up value in those others seasons. There are some other odd outliers, too. Eckersley, for example: the reason he shows up is because most of his value was as a starter, but his seasons as a reliever were so good that they were quite valuable, too. So he shows up, but probably not the way it's meant. So here are the classic "compilers," as the argument is made, who show up in this cursory survey (Player, Rank; Peak WAR, Career WAR, %):

Don Sutton, #6; 52.0, 86.8, .599
Tommy John, #11; 37.4, 59.7, .626
Bert Blyleven, #21; 64.9, 97.0, .669
Pete Rose, #22; 54.1, 80.3, .674
Phil Niekro, #23; 68.7, 101.4, .678
Gaylord Perry, #27; 65.2, 95.5, .683
Lou Whitaker, #28; 45.8, 67.1, .683

So, it actually appears that there may be some validity to this criticism of these players, after all. Particularly, there is a lot of question about John, because his peak was pretty unspectacular. However, I would say that, excepting an extreme case like his, the "compiler" argument doesn't hold water, because most of the greatest players of all-time were compilers, too. I guess I just don't see, "You have a lot in common with Hank Aaron" as being that bad of a criticism, is all.

Thanks to The Baseball Gauge for the data (see above for link.).

Saturday, January 28, 2012

Baseball Hall of Fame Rankings

For a while, I've been envisioning a system to rank players by Hall of Fame worthiness. There are a couple of problems, though. First, I had no access to good data. Second, the market is already flooded with good systems - Bill James Hall of Fame Monitor and Hall of Fame Standards (which rank likelihood of entering the Hall, not merit), Jay Jaffe's JAWS system at Baseball Prospectus, Mike Hoban's CAWS system at Seamheads, and Adam Darowski's wWAR system at Baseball Think Factory. They each use different mathematical models and bases - James uses standard, "newspaper" statistics; Jaffe uses Baseball Prospectus' WARP (Wins Above Replacement Player); Hoban uses Bill James Win Shares; and Darowski uses rWAR (or, as some would call it, bWAR) - the system invented by Rally (Sean Smith) and the most common WAR system, thanks to being hosted at baseball-reference.com. While the systems are different (and I don't really want to get into the nitty-gritty here - go to their sites to find out if you're really interested), they reach largely the same conclusions. So why would I want to do it myself when I don't have the data and there's already good stuff out there?

Well, for one, I have always found value in doing something myself, even if others have already come up with a way to do things. Second of all, these systems all do some things with which I disagree. Third, I found a way to get the data I needed. Over at one of the sites I frequent, The Baseball Gauge, the proprietor, Dan Hirsch, has all of his data free for download. I actually needed a little extra help, but he's super nice about stuff, and we corresponded over e-mail and he helped to give me what I needed. He developed a WAR system over there (Base Runs for offense; Runs Saved, similar to Win Shares, for defense; DIPS 2.0 for pitching), and it's free to use.

So now that I had the data, what was my problem with the other systems. Well, several of them (and others online) use some arbitrary cutoffs - something along the line of [peakWAR+careerWAR]/2. That's the basic formula. Of course, there's really nothing wrong with that. However, how does one define "peak" WAR? 4 seasons? 5 seasons? 7 seasons? 10 seasons? I've seen all of those iterations. So I went to work on the problem. Here are the basic principles to which I stuck through this process.

1.) 10 seasons is the key. Why? Well, it's not arbitrary, for one. The Baseball Hall of Fame requires 10 seasons played for entry. Since that's one of the few rules, it makes sense to me to stick to it.

2.) Big seasons are better than consistency. Imagine two guys with 8 WAR. One of them has 7 WAR one year, 1 the next. The second guy has 4 each year. I prefer the first guy, because with a season that big, you're almost guaranteed to make the playoffs, and have a shot at the World Series. With the second guy, well, lots of guys manage 4 WAR. That may not help the team win. And sure, the first guys may not help at all the second year (with only 1 WAR, he'd be a sub-average player), but, as they say, "Flags fly forever" - in other words, the goal is to win, and you can't take a win a pennant. So the more a player helps to win a pennant, the more it counts. Of course, I'm not counting actual pennants, so I'm going with the seasons that give you the best chance of winning pennants - the biggest seasons.

3.) More different things will give a better estimate than just one thing. I used three different systems, each consisting of two parts. I'll explain.

So, here's the actual system. First, rank the seasons, descending from best WAR to worst. Then...

#1: Add up total WAR. (=A)
#2: Add up total WAR in top ten seasons. (=B)
#3: Add up total WAR, counting the top season 30 times, the next season 29 times, the next season 28 times, the next season 27 times, etc., until you've exhausted all the player's seasons. Then divide by 30. (=C)
#4: Add up total WAR in top ten seasons, counting the top season 10 times, the next season 9 times, the next season 8 times, the next season 7 times, etc., until you've used up all ten seasons. Then divide by 10. (=D)
#5: Add up total WAR, counting the top season 45 times, the next season 44 times, the next season 43 times, the next season 42 times, etc., until you've exhausted all the player's seasons. Then divide by 45. (=E)
#6: Add up total WAR in top ten seasons, counting the top season 15 times, the next season 14 times, the next season 13 times, the next season 12 times, etc., until you reach the tenth season, which will count six times. Then divide by 15. (=F)
#7: You'll now have six numbers, any of which could be the best indicator. So now, we take the harmonic mean of the six numbers:

6/[(1/A)+(1/B)+(1/C)+(1/D)+(1/E)+(1/F)]

Voila!

For example, here are some lists. The top 11 at third base:

schmimi01 Mike Schmidt 94.7 70.9 75.5 41.3 81.9 51.2 63.9
matheed01 Eddie Mathews 88.2 68.9 71.1 40.7 76.8 50.1 61.6
boggswa01 Wade Boggs 76.5 61.9 62.8 38.5 67.3 46.3 55.8
brettge01 George Brett 78.0 59.4 63.1 36.6 68.1 44.2 54.5
jonesch06 Chipper Jones 71.9 52.8 56.7 31.2 61.7 38.4 48.1
bakerfr01 Frank Baker 56.5 54.4 49.6 35.4 51.9 41.7 47.0
santoro01 Ron Santo 56.1 53.4 48.9 34.1 51.3 40.6 46.0
hackst01 Stan Hack 58.6 49.4 48.9 30.8 52.1 37.0 44.0
evansda01 Darrell Evans 59.1 45.5 47.8 28.4 51.6 34.1 41.7
mcgrajo01 John McGraw 48.0 46.9 42.8 32.4 44.5 37.3 41.2
rolensc01 Scott Rolen 53.7 47.7 45.2 28.7 48.0 35.0 41.1

The top 12 at catcher (because I was part of the discussion over at Baseball: Past and Present):

benchjo01 Johnny Bench 75.2 63.5 63.0 39.8 67.1 47.7 56.7
cartega01 Gary Carter 68.7 60.2 57.6 37.0 61.3 44.7 52.5
berrayo01 Yogi Berra 75.2 58.4 60.4 34.6 65.3 42.6 52.3
piazzmi01 Mike Piazza 61.0 54.7 52.2 35.3 55.1 41.8 48.3
dickebi01 Bill Dickey 67.5 52.9 54.9 32.7 59.1 39.4 48.1
cochrmi01 Mickey Cochrane 58.7 53.8 50.0 33.0 52.9 40.0 46.2
fiskca01 Carlton Fisk 67.8 47.4 53.2 29.5 58.1 35.5 44.8
hartnga01 Gabby Hartnett 64.4 47.5 51.3 29.1 55.6 35.2 43.8
rodriiv01 Ivan Rodriguez 65.2 47.0 51.2 28.3 55.9 34.5 43.4
torrejo01 Joe Torre 55.3 46.2 46.2 29.5 49.2 35.1 41.6
simmote01 Ted Simmons 52.5 47.9 45.3 28.6 47.7 35.0 41.0
tenacge01 Gene Tenace 48.0 44.3 41.4 28.5 43.6 33.7 38.6

How about the only 3 DHs I checked, because that's a short list:

molitpa01 Paul Molitor 64.9 49.9 51.8 28.8 56.2 35.9 44.4
martied01 Edgar Martinez 57.9 49.5 48.1 29.2 51.4 36.0 42.9
baineha01 Harold Baines 36.9 28.1 30.0 17.4 32.3 21.0 25.9

One last one. How about top 11 CF:

cobbty01 Ty Cobb 153.2 95.1 112.3 57.3 126.0 69.9 91.4
speaktr01 Tris Speaker 138.5 88.5 103.7 52.3 115.3 64.4 83.9
mantlmi01 Mickey Mantle 120.4 87.8 95.1 53.4 103.6 64.8 81.1
mayswi01 Willie Mays 136.5 85.0 101.0 49.0 112.8 61.0 80.4
dimagjo01 Joe DiMaggio 80.8 69.3 67.6 42.0 72.0 51.1 60.7
griffke02 Ken Griffey 74.4 59.6 60.2 35.5 65.0 43.8 52.9
hamilbi01 Billy Hamilton 70.8 61.2 58.9 36.0 62.9 44.4 52.8
snidedu01 Duke Snider 62.5 54.5 52.8 34.9 56.0 41.4 48.4
wynnji01 Jimmy Wynn 57.4 55.5 50.1 35.3 52.6 42.1 47.4
careyma01 Max Carey 67.7 52.6 54.5 31.3 58.9 38.4 47.2
edmonji01 Jim Edmonds 59.2 52.0 49.7 31.9 52.9 38.6 45.3

Hope that was interesting. I'd love to hear thoughts about this kind of stuff.

As one final thing, I'd just like to thank Dan Hirsch for all his help, and to recommend The Baseball Gauge to anyone out there reading this. It's a great resource.

Tuesday, January 10, 2012

The Baseball Hall of Fame, 2013 Preview

Long time no write!

As of my last writing, Ron Santo was the newest inductee to the Baseball Hall of Fame. Well, on his heels yesterday came the announcement that Barry Larkin would be joining him. Congrats to Barry. My all-time favorite non-Brewers player is going into the Hall! He's really the first player I've ever had any emotional attachment to who got the call to Cooperstown, so I'm pretty excited (although that excitement is just a bit muted by the fact that his election this year was pretty inevitable).

Anyway, what better time to talk about next year's ballot than today? Is it a bit disrespectful to the inductee? Nah - yesterday was his day, and he gets an even bigger day this summer. Rather, I think that, with one of the most loaded ballots in history coming up next year, it's about time we discussed that. So what will next year's ballot look like?

Here are the holdovers from the ballot:

Jack Morris
Jeff Bagwell
Lee Smith
Tim Raines
Alan Trammell
Edgar Martinez
Fred McGriff
Larry Walker
Mark McGwire
Don Mattingly
Dale Murphy
Rafael Palmeiro
Bernie Williams

And here are the (notable/vote-able) guys who are going to appear on the ballot for the first time:

Barry Bonds
Roger Clemens
Mike Piazza
Sammy Sosa
Curt Schilling
Craig Biggio
Kenny Lofton
Upgrade-of-Jack-Morris-with-worse-press . . . I mean, David Wells

Then, the handful of guys who deserve a couple of votes next year:

Julio Franco
Steve Finley
Reggie Sanders
Jeff Cirillo
Shawn Green
Maybe the all-time record holder for most fingers, Antonio Alfonseca

Compared to this year's first-year class, or as I like to call them, "Bernie Williams and the Pips," it's pretty amazing how stacked next year will be. One of the big predictions people are making is that, because Jack Morris garnered 2/3 of the vote this year, his election is inevitable. Well, I could actually see him taking a step back next year because of the loaded ballot. And I'm not really sure that ANYONE will garner election next year. If I had to guess, I would be Morris, Schilling, or Biggio (in no particular order) who would be most likely, because of the steroid stain on the others. And the next year's not much better. It will add:

Greg Maddux
Frank Thomas
Tom Glavine
Jeff Kent
Mike Mussina
Luis Gonzalez
Eric Gagne
Moises Alou
Kenny Rogers

All of those guys should get SOME votes, and I would venture to say that Maddux is automatically a first-ballot guy.

Anyway, what I really want to write about is what my ballot would have been/would be this year, next year, and the year after.

First, my strategy. I think the Hall voters are too stingy. Were I one of them, I'd vote for ten guys every year until there weren't even 10 remotely electable guys left. Sure, I might use one of those ten as a "courtesy" vote, but I'd fill my ballot. This year:

Barry Larkin
Jeff Bagwell
Tim Raines
Alan Trammell
Edgar Martinez
Fred McGriff
Larry Walker
Mark McGwire
Rafael Palmeiro
Brad Radke

My tenth vote could have just as easily gone to Dale Murphy, but I would have shown Bradke some love. Anyway, with Larkin (and Radke) disappearing from the ballot and the aforementioned new class, here is my preliminary ballot for next year:

Barry Bonds
Roger Clemens
Jeff Bagwell
Tim Raines
Alan Trammell
Sammy Sosa
Curt Schilling
Craig Biggio
Edgar Martinez
Mark McGwire

That was really, really tough. And like I said, I don't really expect anyone on next year's ballot to get elected. So that creates my 2014 ballot:

Barry Bonds
Roger Clemens
Greg Maddux
Jeff Bagwell
Frank Thomas
Tim Raines
Craig Biggio
Curt Schilling
Mike Piazza
Mark McGwire

That's right, I left off Alan Trammell, Edgar Martinez, Sammy Sosa, Fred McGriff, Larry Walker, Rafael Palmeiro, Tom Glavine, Mike Mussina, Kenny Lofton, and Dale Murphy (all of whom I see as HOF guys) off the ballot. That's ten guys - a FULL BALLOT. So if no one gets elected next year, which seems reasonable, the backlog becomes pretty much untenable. I don't know that pretty much ANYONE ever gets elected again, because there may always be too many bodies stuck in the doorway, so no one can get through. It will be interesting to see how all of this player out.

Monday, December 12, 2011

The Baseball Hall of Fame

Finals time has a funny way of catching up with me. It makes for little posting here. Oh well. Also, I think I'm supposed to have an obligatory Ryan Braun post, but I'm not going to do that. I'm not emotionally ready, for one. Second, I would like MLB to make its decision before I say anything. That's all I have to say about that for now. On to the topic at hand.

It's that time of year for one of my favorite topics: the Baseball Hall of Fame elections. As anyone who hangs around in the baseball corners of these here internets knows, Ron Santo was recently elected by the Veterans' Committee by an overwhelming vote, with 15/16 voters saying "yes" to the longtime (and now, unfortunately, deceased) Cubs player/fan/announcer (who did play one year for the White Sox).

As for Santo, he was clearly a great player. Whether or not he deserves enshrinement is perhaps a different discussion, but I will say this: using the current statistical standards of the Hall of Fame, Ron Santo easily merits admission. Whether or not that standard is a good or bad one is a topic for a different post (hopefully up later this week). What I'd like to talk about, and what's particularly relevant in the case of Santo, I think, is what it is okay and not okay to consider when making a Hall of Fame vote.

First, let me link to the Hall of Fame's website, where the voting instructions for the BBWAA are listed.

You may notice some interesting things on there. For one, and I did not know this, there is nothing in the guidelines that the player must have been "outstanding" or even "good." There's nothing at all about performance. Interesting on all sides of the Hall of Fame debates, no? There are just the time constraints (been retired for five years, only 15 years on the ballot, 10 year playing career). But really, ethically, what's acceptable to do? That's what I'd like to talk about.

1. Playing ability
Obvious, no? But let's start with the obvious. You need to consider it. If we ignored it, I'm sure things would be a lot different. Perhaps "ability" isn't even the best word: I suppose "results" would do better. And the way we measure these things are statistics. So a player's statistics should count more than anything.

2. Circumstances
Broad category, and I apologize for that, but I can't help it. What I mean here is the basics: era, position, and extenuating circumstances. This helps to flesh out #1. Was his career interrupted by war? Injury? Segregation? Was the schedule shorter? Longer? Are there accurate statistics for this player? Were steroids or other drugs an issue? You may adjust (or not) for any and all of these things as you see fit, but it IS important to keep them in mind.

3. Subjective opinions
Obviously, we can use our own opinions here. Keep in mind, this is the THIRD criterion. That means that the other two trump it. Even so, sometimes it's helpful to check these things. Especially for the players for whom there is little or no data. It can be very beneficial for borderline players. And if we haven't seen a player play, it's nice to read what others said about him. However, keep in mind that this is not the primary way in which we measure a Hall of Fame player.

So, what can we NOT do to measure the Hall of Fame candidacy of a player?

1. Artificially raise the standards for enshrinement
Once standards have been set, it's not fair to artificially raise them. You can't all of a sudden start not electing people who are clearly eligible. However, if standards are deemed too high and need to be lowered, that can absolutely be done.

2. Consider a player's character
This is the one I'm sure I'll get the most flack for. People will say that OF COURSE you have to account for a player's character. But the truth is, we don't know anything about these people. We see them in the little snippets. We hear what others have to say about them. We "learn" things like "he's a racist" or "he's a druggie." But we don't know what led to those things. If it adversely affected his performance or baseball overall, so be it (maybe Pete Rose, or Joe Jackson, or Cap Anson would qualify under either of these). If not, though, it probably should not be considered.

So those are my thoughts. Any opinions out there?