As much as you guys discuss a lot of the benchmark's key limitations, still think you end up letting METR off the hook a bit! Just published a piece criticizing the Long Tasks benchmark myself, albeit in much stronger terms. Think it'll be of interest:
As much as you guys discuss a lot of the benchmark's key limitations, still think you end up letting METR off the hook a bit! Just published a piece criticizing the Long Tasks benchmark myself, albeit in much stronger terms. Think it'll be of interest:
https://arachnemag.substack.com/p/the-metr-graph-is-hot-garbage
Thanks for the link! Looks interesting