Contra Scott Alexander on AI Superforecasting
Oh yeah, and our Tesla backed itself into a concrete pillar. Agonizing about my reading-books-in-the-driver's-seat policy ensues.
(The big news this week is that OpenAI’s GPT-5.6 Sol is out and claims to be about as good as Anthropic’s Claude Fable. Also we got a reprieve on the Fable deadline: you’ve got through Sunday now without paying by the token. I recommend paying Anthropic their $20 just to see for yourself what Fable can do for you this weekend. Today I’m trying both Sol and Fable and my sense is Fable is smarter.1 These aren’t themselves big steps closer to AGI but shooting past the best humans at much of math and coding is a big deal. Recursive self-improvement, when all hell breaks loose, feels close. But perhaps when I get used to these new models, and if we finally see a lull, then I’ll chill out.)
Why I disagree with Scott Alexander on forecasting bots
This isn’t a very good beef because Scott Alexander prudently hedged his predictions. It’s not like he’s confidently asserting that forecasting bots are about to crush the best human forecasters. It’s also not a good beef (for me) because Scott Alexander is right about almost everything. Still, I want to make my case that (a) Scott’s being too credulous about extrapolating the progress of AI forecasting, and (b) he’s giving the “AI as Normal Technology” folks too much credit in the world where they’re right about forecasting having “irreducible error”.
Backing up, remember when people were excited about prediction markets around the turn of the millennium? Well, they were. (It’s pretty weird to me how prediction markets finally went kind of mainstream. And a little depressing how they’ve done so in what’s often a pretty sleazy way, stoking gambling addictions and such.) Anyway, here was Paul Tetlock (son of Philip, of Superforecasting fame) et al in the Wall Street Journal in 2007:
Imagine the president had a crystal ball to predict more accurately the impact of broader prescription coverage on the Medicare budget, the effect of more frequent audits on tax compliance — or even the consequences of a political settlement in Iraq on oil prices. Now, stop imagining: Such crystal balls [prediction markets] are within our grasp.
Some coauthors and I got annoyed by how out of control the hype was (I insist my hyping vs anti-hyping is well calibrated!) and tried to answer the question empirically: how much better are prediction markets exactly? Our paper is called “Prediction Without Markets” and we looked at American football games, baseball games, and Hollywood box office revenue. The answer turned out to be that prediction markets predict better but the difference is pretty razor thin. Just conducting a poll, or setting up a dirt-simple statistical model, is about as good.
What’s going on? My latest take is that most predictions are about as hard as predicting next month’s weather. These are chaotic systems. A sports game between two top teams is roughly a coin flip. You can eke some alpha out by factoring in home team advantage and the barest sliver of additional alpha by scrutinizing the star players’ Twitter feeds or something but not even a superintelligence2 can do much better than predicting the top-seeded team will win with 60% probability or whatever the most dirt-simple heuristic is. And it’s like that for most things we care about predicting.
Which is to say, I think Scott Alexander will turn out to be wrong in his expectation that AI forecasting is about to go significantly superhuman (slightly superhuman, sure).
But my bigger disagreement with Scott is about this being an important test case for the AI as Normal Technology (AINT) thesis. The AINTers think chess and math are the exceptions — that in most domains humans are already close to the ceiling of how good it’s possible for an intelligence to be. (Actually I’m not sure they predicted how good AI would get at math either.) I think it’s the other way around. Chess and math and programming are the canaries in the coal mine about to suffocate us. I.e., AI will eventually get wildly superhuman at most things, just not forecasting in particular.
It’s as if the AINTers were asserting that AI is not on track to beat humans at Candy Land3 and that therefore superintelligence is unlikely to be a big deal.
(Their other concrete test case for AI turning out to be normal technology is persuasion. Fable and Sol are infodumping at me about how AI is already superhuman at persuasion but I’m genuinely not buying it. I’ll tentatively accept this one as a meaningful test case though.)
Your Tesla did what now?
It backed itself into an unmissable yellow-painted concrete pillar while parking yesterday, scraping and denting itself, womp womp. Just surface damage though. Parking has always been a huge weakness for “Full Self-Driving (FSD)” but I presumed till now that actually hitting something would be rare enough to not worry about. But here we are, within the first 5,000 miles and crunch.
Now I’m agonizing about whether I can rationalize this as a parking-specific problem or if I have to rethink my reading-books-in-the-driver’s-seat policy.
I wasn’t in the car myself at the time; my teenage son was. So canonically the perfect scapegoat, but I’d been urging my whole family to always let the car cook, much to their general protesting. And, yes, Tesla blatantly gamified this with an FSD streak feature, which I’m inexorably drawn into. So 1000% my fault. But I have no regrets. If we’d disengaged and saved the car we’d never have been sure if the car would’ve also saved itself at the last second. Science!
So as for Bayesian-updating on this. Well. First, this is absolutely the kind of incident the passenger-seat human safety monitors in the Austin robotaxis (that they had for the first 6 months or so, and even now sometimes) may have been preventing from happening all the time. Unlike incidents at driving speeds, which we’ve recently learned the humans not in the driver’s seat couldn’t prevent. Which is one reason all my number crunching on the public data led me astray in assessing this parking risk. Another reason is that the robotaxis mostly pull over rather than park. Basically, they’re working around their pathetic parking skills by just avoiding attempting it.
That’s good news in one sense: it explains the seeming discrepancy between the robotaxi crash statistics and our family’s abysmal “1 crash in 5k miles”. Chalking the latter up to bad luck would be super sus.
In conclusion (not an actual conclusion, please tell me how cope-addled this sounds), it’s only parking-type incidents that are suppressed in the robotaxi data. The dearth of injuries over those 2 million miles (plus other evidence I’ve been writing about) still means Tesla FSD v14, HW4, in fair weather etc etc is superhumanly safe even without supervision. I’ll just add “car is not trying to park” to my growing list of conditions under which it’s reasonable to ignore what the car’s doing and it’ll be fine.
In any case, next week we’re due to have new robotaxi data so I’ll plan to update my data analysis and reassess. But it’s looking like Tesla is quasi-pausing the robotaxi program, awaiting full self-driving version 15, this time for really realsy reals (FSD15TTFRRR).
Fifty-Two Friday Flashback
A year ago I tackled the argument that superintelligence can’t be dangerous or turn the world upside down because intelligence isn’t real. The thing that’s real is humans’ ability to, as I put it back then, “irrigate deserts, change global temperature, eradicate diseases, engineer new diseases, drive a car on the freaking moon, you name it.” If an AI lives on the internet you might expect it to be harder to do things like that. I don’t think so. It can hire human labor just like any business. Or build better robots.
I like these diagrams from LessWrong/Twitter, that I’ve pointed to before, showing two ways AI intelligence/capability might increase:
See also the “Ord Cloud” analogy.
AI being good at some things and bad at others doesn’t mean it won’t eventually crush us at everything. Again, intelligence being multidimensional doesn’t save us.
In other news from a year ago: MechaHitler. Fun times. This week even Elon Musk admitted that Fable is smarter than the latest Grok (xAI/SpaceX’s LLM). Last year I also added:
In other news, Grok 4 just came out. I’ll keep juuuust enough of an eye on it to be able to let you know if it ends up pushing the frontier in any non-nazi ways.
I can confirm that Grok hasn’t been especially close to the frontier in a long time.
Next, here’s a fun one to look back on:
Some doubt’s been cast on whether AI coding assistance actually saves time on net. It definitely saves me time but maybe if I sucked less it wouldn’t? Another (less embarrassing for me) theory is that there’s a learning curve to using these tools well, and the study included a lot of people who’d never tried them before. Also, to be fair, I do sometimes waste a stupid amount of time arguing with AI instead of rationally giving up as soon as it becomes apparent it’s not smart enough to do something I thought it was smart enough to do.
A year later, all doubt is long gone. This isn’t just me being bad at coding anymore. In some ways I am, in other ways not — my favorite is Project Euler problems. Last weekend I wrote a Sudoku solver (by hand!). But whatever coding you’re best at, if Fable isn’t better than you, I think its successor will be. Even if you’re the best in the world. A programmer claiming to be slowed down by AI coding assistance would be like a cyclist saying that using these janky newfangled auto-mobiles only slow them down. I mean, maybe you’ve got a janky enough car that that’s true. You’re definitely more energy-efficient without the car. And maybe you love biking and hate driving. That’s super fair. But if you just need to get somewhere and speed matters, the car wins. And it’s not going to be close, once the janks are ironed out.
Random Roundup
Anthropic’s new J-Space research seems to be a breakthrough in looking deep into an AI model’s brain (not just its chain-of-thought scratchpad). They even have demos where you can try things like asking the LLM how many legs the critter that spins webs has, wait for “spider” to light up in its brain, reach in and change that “spider” to “ant”, and, lo, it answers “6”. Spooky. Btw, the J in J-Space is for Jacobian, a generalization of the derivative in calculus, and is part of the math that makes all this work.
Disclosure: I don’t get it yet myself.
The AI Futures Project has a new write-up called AI 2040: Plan A (they also have a blog post introducing it, and Scott Alexander has a long summary). Recall the original AI Futures Project write-up, AI 2027, that I’ve often mentioned. I have to re-explain this every time that despite the ill-advised title, AI 2027 predicted a wide range of possible futures with huge uncertainty about time to AGI. They picked their modal (most likely) year and described a slightly fictionalized scenario where fully superhuman coding (no more software engineers) happens by then, in detail, for concreteness. It seems to be proving prescient in some ways so far, like predicting something like the drama that played out with the US government blocking the release of Fable. And of course the wild advances in AI coding and AI spending — not at all obvious in early 2025. (The pooh-poohers expected everything to implode in a pile of naked emperors by now.)
One more news item: OpenAI launched GPT-Live. The demo makes it seem like they’ve finally cracked natural back-and-forth verbal conversation with AI. But every demo (going back decades!) has seemed like that and it’s always lies. Sure enough, I tried GPT-Live (so you don’t have to) and it’s as excruciating as ever. You’re welcome.
PS, by popular demand, here’s what the actual damage looks like:
As a random example, see the spiky/blobby diagrams at the end of today’s newsletter? I asked Sol and Fable, “can you find the ai spikiness diagram from an old agi friday?” Only Fable could, undaunted by me calling them, back in December, “jaggedness diagrams” without ever saying “spiky” or “spikiness”. PS: Gemini also got it.
Or, fine, maybe all bets are off for an actual superintelligence. I’m talking about what we can expect in the next couple years, which I’m presuming is still pre-AGI.
Candy Land has no skill component so there's no such thing as anyone being better than anyone else at it.







I am vowing not to use GPT-Live except for specific purposes because the conversations I was having with it were lulling my mind into breaking its hard earned habits of not interrupting the other speaker. I am proud of generally not interrupting other people in conversations but something about live broke that habit for me temporarily and I saw it manifest slightly in real life as well.
I was really shocked by the micro improvements.
A) I can have a long, and winding thought without it interrupting me if I pause, with the model waiting for me to finish my thought. Previously the model would interrupt with any silence.
B) Live can laugh *alongside* you as opposed to *after* you in the right context. It’s extremely scary and made me step back for a second.
I find the model highly unpleasant to talk to when having normal conversations as I find it uncanny, and uncomfortable to talk to with someone who is so patient to hear whatever rambles you have. I am used to being the listener in the conversation so having a [virtual] conversation partner that is so patient and unopinionated…
I may still use live for French practice as it is waow bananas at naturalistic conversation. Neither GPT-Live nor the previous model understood vietnamese, oddly.
If you want your mind blown, try using Live instead of Google Translate next time you need to chat to a Spanish French or Portuguese speaker. Ask it to be a two way interpreter.
When I worked a stint at best buy I relied on the previous voice model daily to act as my interperter and it worked way better than google translate, but would outright fail randomly maybe a third of the time to even enter an interperting mindset when prompted.
I'm curious about that parking spot. Given the position of the red car, and the yellow post, was there even room for your car? Or does the FSD choose the parking spot?