What an agent harness does
A coding agent needs a way to inspect files, run commands, observe failures, and continue working. The harness supplies the surrounding execution loop: context, tools, state, and rules for deciding what happens next.
OpenAI describes the Codex loop as planning, editing, running tools, observing results, repairing failures, and updating status. Its account of a long-running coding experiment also emphasizes durable project notes and checks at each milestone. The model operates within that system; a good prompt is only one part of it. Source: OpenAI's long-horizon Codex experiment.
Design around an observable result
For a portfolio application, consider the task “make social links editable.” A useful acceptance condition is concrete: an administrator changes an address, saves it, reloads the editor, and sees the same link on the public page.
That gives the harness meaningful feedback. A compiler check verifies types. A persistence check verifies storage. A browser check verifies the interaction. Each catches a different failure, so passing one should not automatically mark the entire task complete.
A practical starting design has four pieces:
- A short task description with the intended behavior and constraints.
- Tools with explicit inputs, bounded execution, and useful error messages.
- A progress record containing decisions, completed checks, and remaining work.
- A stopping condition tied to the actual user outcome.
These are engineering suggestions, rather than a required framework or a promise about any model's reliability.
Make recovery deliberate
Imagine a content save failing because the remote branch changed. Repeating the same write indefinitely would consume time without addressing the cause. A better application workflow reloads the current state, checks whether the change is still valid, and retries within a limit.
The same principle applies to an agent. Record the failure, choose a recovery step, and verify the result. Keep externally visible actions behind the authorization the task actually provides.
Measure the whole workflow
My suggested evaluation is a small set of representative tasks, including missing context, unavailable tools, and partial failures. Track successful outcomes, recovery attempts, elapsed time, and whether changes stayed within scope.
Changing the model can help, but changing how the agent receives feedback can also change the result. Treat both as parts of the system you evaluate.