End goal: a leaderboard of harness performance (multiple axis) on a set of diverse real world tasks[1], grouped by underlying models and reasoning efforts. Anyone can contribute results.
The task criteria, measurements, underlying framework et al can be decided by a group rather than a single person.
If there is sufficient interest, I will create a discord.
Disclosure: I am the maintainer of a coding agent called Dirac (https://github.com/dirac-run/dirac) so I will not influence what the final benchmark should look like to avoid any conflict of interest. I just want to make this happen.
[1] Diverse real world tasks meaning sufficiently complex tasks that the contributors have encountered, preferably from an opensource repo.
So you have any specific ideas?
You can hmu at iam@thechris.in
> We can do a mix of general use (as in user stories) plus a few academic benchmarks.
Yup sounds about right. Generally speaking, higher the distinct contributors, more likely it is to capture the distribution of real-life usefulness.
> So you have any specific ideas?
Only that the problems that get picked should be easy to evaluate in isolation and should test the harness capability rather than model's knowledge/capability.
Then we need a provenance for model inference, generalized. This should be interesting. We would be trying to deterministically generalize a baseline "can do this" for models... Maybe categorize by parameter class.