I’m not quite sure why the harness itself needs to be small. Isn’t the system prompt and management of system prompt the bit you want lightweight?
bitbasher 2 hours ago [-]
Here I was thinking 1MB still feels way too large.
userbinator 20 hours ago [-]
"Why does it need to be large?" should be the question instead.
selcuka 19 hours ago [-]
A harness needs to be good. Being large is annoying, sure, but it's generally not an issue. Unless this one has feature parity with the "large" ones (hint: it doesn't), I don't see any reason to use it only because it's small.
paoloanzn 12 hours ago [-]
I ported almost every key feature already tbh. Of course there will always be a compromise, but there is even still margin to add more. Tools, skills, context compaction, it’s all already implemented.
selcuka 12 hours ago [-]
I haven't actually tried it, my comment was based on the "Known bugs and issues" section in the readme (e.g. no MCP support yet). Apologies if that section is out of date. If the functionality is on par with the mainstream ones, then sure, small and lightweight is preferable.
paoloanzn 4 hours ago [-]
Actually the MCP support is really one of the only "key" things we had not brought over yet, but please feel free to open an issue or a PR!
cyanydeez 18 hours ago [-]
ultimately, the harness is just going to be another model specifically engineered to be deterministic, like ballbearings in a wheel minimizing friction.
K0IN 19 hours ago [-]
I think it's nice to have options, but I also don't see the appeal in "sub x mb" (except as codegolf or just for the hacky part, for day to day this does not matter) especially if features are missing or other compromises have been made just for the sake of remaining small.
paoloanzn 12 hours ago [-]
I’m mainly looking also at running a cli agent inside a small device, with few resources.
embedding-shape 16 hours ago [-]
Probably you want all of them to be as simple and small as possible, especially if you want to use it to work on itself. Smaller codebases are (generally) easier to work with.
fractorial 2 hours ago [-]
For context, I've rolled my own research harness in Go; I've been trying very hard to keep the non-test LOC around ~30k lines. To me, I consider this large, but necessary, especially after meticulously crafting red diffs while adding features.
I do use it to work on itself. I agree smaller is likely easier to work with; however, that's still, at least partly, disjoint from binary size.
adithyassekhar 15 hours ago [-]
Sometimes smaller software has the harder codebase.
embedding-shape 8 hours ago [-]
Damn, you really got me there, nothing in my comment could have possibly have guarded against that not every smaller codebase is easier to work with. If only there was a way in English language to indicate side-thoughts, optionals and asides, but oh well.
paoloanzn 12 hours ago [-]
What about running agents on embedded devices with small resources? maybe even multi agents flows? good luck trying to do that with codex, let alone claude code
orliesaurus 18 hours ago [-]
i love the idea, but you also need to maintain it long term: in today's world what is the best practice to do that... do you spin up an agent anytime codex-cli updates and mirror the updates in c++ and push a new release? or idk... set up an automated process that does this on a worker on CF in a sandbox when npm updates @openai/codex?
paoloanzn 12 hours ago [-]
there is already a wrapping script when you install it, every time i publish a new release it will detect it a prompt it to update (if you want to).
I do use it to work on itself. I agree smaller is likely easier to work with; however, that's still, at least partly, disjoint from binary size.