AI safety researcher
Thomas Kwaš¹
Itās plausible to me that it inspires ambitious behavior more than unhealthy behavior. Most people do not have obsessive personalities, and ā¬1M is a reasonable amount for a dedicated person to donate over a few years.
84 hour weeks are also common in industries like investment banking. If someone said they were giving up all distractions, working 84h weeks as an investment banking VP and donating 70% of their salary for 3 years, we would just call them dedicated. So I think the only caveats that are needed are to ensure the path to ā¬1M is realistic and to avoid burnout.
I started writing bc I felt obligated to respond but only continued because the worksheet limit thing was funny. It wouldnāt be funny the second time, so commenting on ārandom slop about METRā probably wonāt eat infinite time. Unless it were much more prevalent I guess.
Apparently the Forum policy allows this and posts are supposed to be automatically flagged, but I donāt see a flag on this post. Agree it looks like Fable.
Most of these criticisms are not new; for an organized writeup of the most important known issues in the original time horizon paper see my blog post from January that OP linked in the conclusion. I do have some comments on this postās methodology.
Task suite (HCALC, n=23). Human-Calibrated Arithmetic and Ledger Computations: quantitative tasks screened for objective, automatic scorability ā grocery receipts, payroll, a 360-payment amortization schedule, descriptive statistics on 500 blood-pressure readings, a 200Ć200 matrix inversion, and two Monte Carlo simulations. Screening for automatic scorability conveniently restricts the suite to things Excel can attempt.
Excel finished the retirement simulation in 51 seconds, a 135,294Ć speedup
The super long tasks are not economically valuable, because no human would do 200x200 matrix inversion or Monte Carlo simulations by hand in 1985, so theyāre not really worth tens of weeks of human wages. The baseline for these should probably be a human C or Fortran programmer with access to a reasonably fast computer, probably an hour or two rather than weeks.
It is also not really true that automatic scoreability restricts the suite to things Excel is good at. E.g. Excel cannot do reasoning questions or computer use, but it does support scripting which is not included here.
The agent. Excel cannot type, so it is scaffolded with a human peripheral who keys in values at a measured 0.3 seconds per entry (n=1, mildly caffeinated) and contributes no cognition.
If the point is that the scaffold heavily affects the intelligence of the model, METR tested this in February and found that Codex and Claude Code scaffolds donāt outperform the standard scaffold we use. Generally the bigger issue has been models not using scaffolds properly than exactly how much optimization goes into the scaffold.
80%-time horizon of frontier systems, 1985ā2026
The graph lists Microsoft Excel 1.0, but the benchmarks must have been run on a modern version of Excel. The worksheet limit of Excel 1.0 was only 16,384 rows rather than 1,048,576, plus it lacked many of the functions of modern Excel, so it would presumably fail more tasks.
The logistic does the work. Two parameters, fit through mostly-ones and a few zeros. Given only my aggregate success rate and the task-length distribution, you could recover the horizon without knowing which tasks Excel passed.
Shashwat Goel showed you can reconstruct the entire log-linear trend from aggregate accuracy plus the task-length distribution with a fixed slope ā the individual task outcomes barely matter.
Itās true that time horizon is highly correlated with success rate if you know the task length distribution, and I used this shortcut in my follow-up last year, for all the benchmarks where we didnāt have individual question data. IMO itās not a major flaw because
is known to differ wildly by task distribution and you canāt estimate it just from aggregate accuracy; there could be some weird distribution on which even isnāt enough.If a capability can sit 36 orders of magnitude above trend for four decades without anyone noticing, the correct response to any capability chart is fear, and the correct response to the absence of a capability chart is more fear.
This violates conservation of expected evidence. It canāt be the case that capability chart and no capability chart are both evidence of dangerous capabilities, and itās not healthy or productive to feel fear regardless of the evidence.
Finally, this post seems almost entirely AI written, probably by Claude 4.8 or Fable 5. Pangram says itās 100% AI written. This created a bunch of minor issues in the writing.
This was actually a deliberate choice. People are more agreeable on EA Forum than LW, so the simplest model that fits the data is that everyone who agrees with a comment on EAF will disagree on LW. The button placement just facilitates this!
It seems possible to vibe code a chrome extension to do this in under 2 hours, maybe under an hour if you have codex or claude code set up already.
A counterfactual donation isnāt just my act being contingent on your act, it means AMF getting a net extra $50 is contingent on your act. So if part of the $50 were a donation match that would be filled regardless, or some (especially low op cost) funder would make up the difference, it wouldnāt fully count.
Counterfactual is also used to mean the second choice opportunity that defines an opportunity cost. āShould I take job A? Well my counterfactual is job B, where I would do X...ā
My guess is something like: Many organizations have quarterly caps on the number of false claims published. Their employees often want to make false claims, but towards the end of the quarter theyāre at the cap, so they delay the post to the first day of the next quarter.
Okay, but why only April 1? Well, on Jan 1 everyone is on holiday, and on July 1 everyone is out enjoying the good weather. Oct 1 coincides with national holidays in populous countries like China and Nigeria, and in the US people are hung over from fiscal New Yearās Eve. So we only really see the effect on April 1.
I would strongly predict that a false claims spike also happens in places with bad weather on July 1. Unfortunately, most places are in the Northern Hemisphere where itās warm, and Australia has good weather all year, so I think this is only testable when it snows in New Zealand.
Iād love to sign up, but due to adverse selection concerns Iād prefer to be matched with an EA picked uniformly at random (whether they signed up or not). Is this possible?
what prompt did you use?
On a global scale I agree. My point is more that due to the salary standards in the industry, Eliezer isnāt necessarily out of line in drawing $600k, and itās probably not much more than he could earn elsewhere; therefore the financial incentive is fairly weak compared to that of Mechanize or other AI capabilities companies.
Being really good at your job is a good way to achieve impact in general, because your āimpact above replacementā is what counts. If a replacement level employee who is barely worth hiring has productivity 100, and the average productivity is 150, the average employee will get 50 impact above replacement. If you do your job 1.67x better than average (250 productivity), you earn 150 impact above replacement, which is triple the average.
I strongly disagree with a couple of claims:
MIRIās business model relies on the opposite narrative. MIRI pays Eliezer Yudkowsky $600,000 a year. It pays Nate Soares $235,000 a year. If they suddenly said that the risk of human extinction from AGI or superintelligence is extremely low, in all likelihood that money would dry up and Yudkowsky and Soares would be out of a job.
[...] The kind of work MIRI is doing and the kind of experience Yudkowsky and Soares have isnāt really transferable to anything else.
$235K is not very much money [edit: in the context of the AI industry]. I made close to Nateās salary as basically an unproductive intern at MIRI. $600K is also not much money. A Preparedness researcher at OpenAI has a starting salary of $310K ā $460K plus probably another $500K in equity. As for nonprofit salaries, METRās salary range goes up to $450K just for a āseniorā level RE/āRS, and I think itās reasonable for nonprofits to pay someone with 20 years of experience, who might be more like a principal RS, $600K or more.
In contrast, if Mechanize succeeds, Matthew Barnett will probably be a billionaire.
If Yudkowsky said extinction risks were low and wanted to focus on some finer aspect of alignment, e.g. ensuring that AIs respect human rights a million years from now, donors who shared their worldview would probably keep donating. Indeed, this might increase donations to MIRI because it would be closer to mainstream beliefs.
MIRIās work seems very transferable to other risks from AI, which governments and companies both have an interest in preventing. Yudkowsky and Soares have a somewhat weird skillset and I disagree with some of their research style but itās plausible to me they could still work productively in a mathy theoretical role in either capabilities or safety.
However, things I agree with
If the Mechanize co-founders wanted to focus on safety rather than capabilities, they could.
the Mechanize co-founders decided to start the company after forming their views on AI safety.
The Yudkowsky/āSoares/āMIRI argument about AI alignment is specifically that an AGIās goals and motivations are highly likely to be completely alien from human goals and motivations in a way thatās highly existentially dangerous.
Is there a formula for the pledge somewhere? I couldnāt find one.
See the gpt-5 report. āWorking lower boundā is maybe too strong; maybe itās more accurate to describe it as an initial guess at a warning threshold for rogue replication and 10x uplift (if we can even measure time horizons that long). I donāt know what the exact reasoning behind 40 hours was, but one fact is that humans canāt really start viable companies using plans that only take a ~week of work. IMO if AIs could do the equivalent with only a 40 human hour time horizon and continuously evade detection, theyād need to use their own advantages and have made up many current disadvantages relative to humans (like being bad at adversarial and multi-agent settings).
A slidĀing scale for donaĀtion percentage
What scale is the METR benchmark on? I see a line that āScores are normalized such that 100% represents a 50% success rate on tasks requiring 8 human-expert hours.ā, but is the 0% point on the scale 0 hours?
METR does not think that 8 human hours is sufficient autonomy for takeover; in fact 40 hours is our working lower bound.
What if we decide that the Amazon rainforest has a negative WAW sign? Would you be in favor of completely replacing it with a parking lot, if doing so could be done without undue suffering of the animals that already exist there?
Definitely not completely replacing because biodiversity has diminishing returns to land. If we pave the whole Amazon weāll probably extinct entire families (not to mention we probably cause ecological crises elsewhere and disrupt ecosystem services etc), whereas on the margin weāll only extinct species endemic to the deforested regions.
If the research on WAW comes out super negative I could imagine it being OK to replace half the Amazon with higher-welfare ecosystems now, and work on replacing the rest when some crazy AI tech allows all changes to be fully reversible. But the moral parliament would probably still not be happy about this. Eg killing is probably bad, and there is no feasible way to destroy half the Amazon in the near term without killing most of the animals in it.
Itās plausible to me that biodiversity is valuable, but with AGI on the horizon it seems a lot cheaper in expectation to do more out-there interventions, like influencing AI companies to care about biodiversity (alongside wild animal welfare), recording the DNA of undiscovered rainforest species about to go extinct, and buying the cheapest land possible (middle of Siberia or Australian desert, not productive farmland). Then when the technology is available in a few decades and weāre better at constructing stable ecosystems de novo, we can terraform the deserts into highly biodiverse nature preserves. Another advantage of this is that weāll know more about animal welfareāas it stands now the sign of habitat preservation is pretty unclear.
The hummingbird should clearly have solved decision theory, then acausally traded with infinitely powerful beings from elsewhere in the multiverse who could create entire fireproof universes.