Without benchmarks and/or a whole suite of non-cherrypicked examples, this means...

softservo · 2026-05-03T22:36:21 1777847781

Working on benchmarks at the moment! Always open to feedback / PRs.

carterschonwald · 2026-05-04T00:03:58 1777853038

im def working on benchmarks for how my own general harness improves task performance vs same model in a commodity setup. its hard to do!

i will say that my current harness: https://github.com/cartazio/oh-punkin-pi is a testbed for a bunch of 2nd gen harness tech, largely optimized for reasoning llms only. the next one after this harness is gonna be epicccc