Contributed by the Risk Threshold Forecasting community.

When will an 8 hour, 80% reliability time horizon be achieved on METR’s Autonomy Tasks by a GPT-4.5 scale model by OpenAI?

Current estimate
Top Key Factors