Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

It is impressive that it (almost) saturates ARC-AGI-3, https://x.com/PrimeIntellect/status/2085087000764568010.

I am curious - how does it fare for other benchmarks, or everyday programming?



PrimeIntelect is not on official ARC-AGI-3 leaderboard: https://arcprize.org/leaderboard


Good to know!

Is it that it wasn't accepted yet, or are there issues with how it was run?


It’s a self-improving harness, and ARC-AGI-3 is explicitly a few-shot benchmark. It’s likely that it gave itself more than the maximum number of tries to learn the games, or even hardcoded the answers.

There’s a lot of improvement to be had from the benchmark harnesses, but sometimes, like with ARC-AGI-3, the limitations are intentional.


leaderboard likely has results from "semi-private" dataset, and graph above likely from public dataset, so it can be easily overfit.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: