← All writing

Objectives: A Lightweight Way to Keep Agents on Track for Long-Running Tasks

A prompt and a few files that keep a coding agent on track for long-running tasks, without building orchestration.

One of the biggest questions in AI right now is long-running agents, i.e., the ability to give an agent a lofty task and have it run without input for hours, days, or weeks until it accomplishes what you asked it to do.

The hard part is keeping it on track. It has to remember what it already proved across compactions and restarts, and it has to know when it's allowed to stop.

I ran into this in May 2026 on a public coding challenge:

build the best image-grouping algorithm for real estate HDR photos
$25,000 for first place.
I knew almost nothing about image grouping or vision models.

I entered for the money and to see how far I could push the models at the time (GPT-5.5).

Discovering Goal Mode

My first attempt was a multi-agent research harness.

Each round went like this:

A planner picked three hypotheses to try.
Three workers each built one in its own git worktree.
Reviewers scored each candidate and looked at the images it got wrong.
A triage agent compared the results, and research agents dug into what failed.
A synthesis agent picked the next three hypotheses, and the loop ran again until the score cleared the target.
My Research Harness
My Research Harness

Goal mode within Codex had just come out, and I wanted to test how good a simple goal could be against my awesome custom harness.

I ran a simple goal "Get 100% accuracy on the datasets"
It cooked.
It cleared 99% in under two hours while my harness sat around 92%

I was extremely humbled, and I decided to see how far I could push goal mode.

I don't think the harness was a bad idea.
Built out completely, it would probably surpass the goal mode / objective
The time commitment is much higher though
Every piece needed tuning: the planner, the workers, the reviewers, the research loop, and the handoffs between them.
The objective got me most of the way there with minimal setup

Constraints of Vanilla Goal Mode

Goal mode got me most of the way, but pushing further raised a bigger question:

What does a long-running agent need to keep doing good work for days?

Four things kept coming up.

Explicit Stopping Criteria

A vague target lets the agent stop at the first easy win.

With several datasets, "get the accuracy higher" is especially loose.
The agent raises the score on one dataset, sees a bigger number, and calls it done.
Mine also overfit.
The public sets were small.
The largest had 5,041 images, against 266,376 in the full training set.
Its first run on the hidden test set scored 92.7 percent, far below the 98 percent I needed to be competitive.

What worked was spelling out every metric on every dataset.

1
2
3
4
random10: mean >= 99.40%, p10 >= 99.25%, worst >= 99.05%
random20: mean >= 99.35%, p10 >= 99.15%, worst >= 99.00%
random40: mean >= 99.25%, median >= 99.20%, worst >= 98.95%
random80: mean >= 99.15%, median >= 99.10%, worst >= 98.90%
Done meant every dataset clearing every gate.
One good number on one dataset no longer counted, so the agent kept going.

How We Want It to Work

The biggest unlock with agents is their superhuman ability to write complex low-level code and then run a huge number of experiments.

The agent was writing complex matrix-backed statistical analysis files of nearly 6,000 lines.
Then spawning a sub-agent to review the results.

It worked, and I wanted that process written down:

Run experiments with MLX on my GPU.
Set them up this way for this project.
Review results like this.
The more the agent knows up front, the more experiments it runs, and the faster it runs them.

Every Experiment It Has Run

One evening of goal mode left 38 experiment scripts at the top of my repo.

They sat next to the handful of files that mattered.
It marked itself done and never cleaned up.
Nothing said which script belonged to which experiment, or which results were current.
goal-test/
+context/
+inputs/
outputs/# 64 run outputs
+Dockerfile
+goal.md
+requirements.txt
+solution.py# the actual solution
-structural_sweep.py
-targeted_geometry_eval.py
-threshold_sweep.py
-visual_rule_eval.py
-… 34 more scratch .py files

Where It Is Right Now

Long runs compact and restart. Each time, the agent needs to know:

The completion criteria.
How to run experiments properly.
What it already proved.
What's still running.
What comes next.

Without that, it repeats work and drifts.

The catch is getting all of this to the agent through goal mode.

The goal prompt is static. It stays in context through every compaction.
Everything else gets compacted.
Run details, results, and decisions get summarized.
On a run that lasts days, I can't trust the summary to keep everything I care about.
So the goal has to carry it all, and it's capped at 4,000 characters.
There's far more context than that.

Enter Objectives

An objective is a way to get all of that context to the agent, simply and efficiently, using just natural language and file system structure.

The goal prompt stays short and points to files.
The files hold everything else: the full stopping criteria, how to work, the experiment record, and the current state.
After every compaction, the goal tells the agent to reread them.

The goal prompt is the key piece.

Every objective's goal has the same shape:

xml
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
<goal>
- What this objective must achieve.
</goal>

<context_refresh>
- Reread objectives/<slug>/goal.md.
- Reread objectives/<slug>/current_state.md.
- Reread the relevant objectives/<slug>/context/*.md files.
</context_refresh>

<working_strategy>
- The approach and the order of work.
</working_strategy>

<success_metrics>
- Observable signs of progress.
</success_metrics>

<non_goals>
- What this objective must not expand into.
</non_goals>

<completion_criteria>
- What must be true before the objective is done.
</completion_criteria>

The context_refresh block does the heavy lifting. Each constraint gets a home:

Stopping criteria live in goal.md.
Every metric on every dataset, so the agent knows exactly when it's done.
Short enough to fit under the 4,000-character limit.
How it should work lives in context/.
The plan, the constraints, and project how-tos, like running experiments with MLX.
No length limit, and the agent rereads it after every compaction.
Every experiment lives in the objective's workspace.
Scripts, reports, and run outputs all land inside the objective.
The repo stays clean, and every file traces back to the objective that made it.
Where it is right now lives in current_state.md.
What's proven, what's running, and what's next.
A restarted agent picks up where it left off, and I read the same file to check on a day-long run.
objectives/
example/
artifacts/# reports and run outputs
context/# how it should work
scripts/# every experiment script
current_state.md# where it is right now
goal.md# stopping criteria and the context_refresh list

The loop has no controller of its own. Every time the agent starts or compacts, it rereads the objective, and the stopping criteria decide whether the next step is another experiment or a stop.

Pros and Cons

Objectives aren't a replacement for a full orchestration system. They're the fast way to get one long task running.

Pros
Very easy to set up
A prompt and a few files, with no service to build.
Built for long runs.
The agent stays on track across compactions and restarts.
Generalizable
It works for most tasks you can describe with clear stopping criteria.
Cons
Not optimized.
You can't tune individual agents to bring cost down or make the run more efficient.
One agent, one path.
There's no explicit fan-out to parallel candidates or separate research agents like the harness had.
You can give instructions to do so, but there's no guarantee it will parallelize.

If you need a general task to run for a long time with a good chance of working correctly, objectives are a good fit. If you need it cheap and efficient at scale, build the orchestration.

Try It

The objectives docs cover setup, the file format, and how to resume a run.

The objectives repository has the skill and its source.