← All writing

Can AI Make You World Class at Something You Know Nothing About?

How I built a vision model that outperformed everything while not understanding how a single line of the code worked.

The internet is full of impressive AI demos, but the real question is:

How do you get immense value from these systems?
How do you get more than human capability out of them, or push into a hard field you know nothing about?

I started without the skills this problem needed. I had never built a vision model, and I did not know how to make code fast on a CPU or GPU.

Using AI (Codex), I built a grouping algorithm at a world-class level.

These are the techniques that got me there.

Ask for the method, not the answer
I knew to worry about overfitting, but not how to split a dataset to prevent it, so I asked.
Turn loose goals into hard bounds
Agents quit at the first easy win until each run had explicit metric gates.
Split optimizing from implementing
One agent proposed speedups, and a second built them with byte-identical output, which cut runtime by 80 to 90 percent.
Point a lot of compute at the problem
Long-running agents ran thousands of experiments on one machine, made cheap by caching, a GPU port, and proxy screening.
Stay in the loop on the data
AI tagged 266,376 images, and my review tools turned my calls into corrections.

The problem came from a public coding challenge.

It asked for the best algorithm to group real estate HDR bracket photos.
First place paid $25,000.
I did not win, but I kept pushing the agents after the challenge ended.

Understanding the Problem

A real estate photographer shoots each camera angle several times through the use of a macro to capture exposures from very dark to very bright.

Those frames are a bracket set, and they get merged into one HDR photo.

Every picture you see online actually is too good to be true

What Goes Together

Frames belong together when the photographer shot them as one bracket set from one camera angle.

Within a set, the furniture, windows, and fixtures stay in the same place, and only the brightness changes.
A second angle of the same room is a separate group, even though it shows the same table and chairs.

Computational Constraints

The container had no GPU and no internet
Run had to complete in 60 minutes on 16 vCPUs and 32 GB of RAM
Cannot use external models and such to parallelize
Grouping had to come from the pixels, through similarity, tone, and geometry.
EXIF metadata, filename parsing, etc. were not allowed to be used
Every image resized to a long edge of 1,024 pixels.
You cannot get a signal from different camera resolutions

The Hard Cases

Getting the general grouping correct is relatively easy, the hard part is then separating groups that are nearly identical. This is done so that when images are merged / processed by AI, it does not induce weird artifacts.

State Change

Where the angle is the same but something in image state change
Example
An entryway with the door closed and again with it open, so these are two bracket sets.

Slight Angle Change

When scene stays the same, but the photographer moved slightly
Example
The camera moved slightly between these two kitchen sets, which shows at the counter edge in the lower left.

Same Floor Plan, Cosmetic Difference

Two houses with the same floor plan produce rooms that look nearly identical and must stay apart.
These bathrooms match in layout, and the fixtures give them away, chrome in one house and black in the other.

Building the Solution

Starting From Zero

I did not know how classic computer vision worked, so the first thing I did was ask the AI.

I asked how image grouping works and how to break the problem into stages.
From there, I set the boundaries for each stage, and the agents wrote and ran the experiments inside them.
My first setup was a multi-agent research harness, and a plain Codex goal beat it.
My objectives post tells that part.

The work turned out bigger than I expected, and I ran out of time before the challenge closed.

My last entry scored about 98 percent, and the winner scored about 98.6 percent
I placed seventh, which won no prize money.
The organizer never paid out the prizes, citing a flawed dataset.
From what I found later, the dataset was extremely flawed.

My final attempt achieved a ~99.4% accuracy across the cleaned full 260K image dataset

Building a Dataset I Couldn't Overfit

From research in college, I knew to worry about overtraining, but I did not know how to build the splits that prevent it.

So I asked the AI how to split a large dataset into tuning pieces and holdout sets that would catch overtraining.
The answer was a ladder of shards, samples of the dataset that keep every labeled bracket set whole, with a job for each size.
ShardsJobRule in the goals
10K and 20K
Smoke tests and pruning
Sweep these cheap cached shards first
40K
Main optimization loop
Primary comparison
80K
Scale confirmation
Required before any promotion
Stress and control sets
Hard cases, replayed failures, neutral samples
Can block a promotion, never make one

Designing the Pipeline

I asked the AI how to structure the grouping, and it laid out a chain of phases that each hand one table to the next.

Rebuilt Grouping Phase Mapsandbox

Each phase does one job, and none of them uses a neural net.

Phase 1
Turns each image into a row of hand-built image statistics
Phase 2
Proposes candidate pairs, the pairs of images worth comparing.
It bounds the work, so no later phase compares every image with every other image.
Phase 3
measures each candidate pair, and stacked XGBoost models turn those measurements into one score per pair.
Phase 4-7
Accepts every pair whose score clears a threshold and joins the accepted pairs into groups with union-find.
Union-find is a data structure that tracks which items share a set.
Each accepted pair, called an edge, merges the sets of its two images.
Phase 8
Packages prediction into output class row

Optimizing the Accuracy Algorithm

The algorithm had to group the photos correctly, and it had to run fast enough for the challenge limits.

What to Optimize at Each Phase

Obviously accuracy is the key metric to optimize, but since my algorithm has multiple independent phases, I did not know what metrics to optimize for each phase, so I asked AI.

It determine key metrics, how they relate to each other and what would be considered amazing bounds if success was achieved

A loose goal let the agents stop early.

Given a goal like "find an improvement," an agent would find a simple win and stop.
So each run got explicit gates, with several metrics above set thresholds across several datasets.
The run kept going until it cleared every gate or I stopped it.
The gates lived in objectives, the goal, context, and state files I wrote up in Objectives: A Lightweight Way to Keep Agents on Track for Long-Running Tas

Phase 2: Candidate pairs

The trade was coverage against cost.
Rows per image fell from 106.0 to 15.0 while maintaining near 0 fragmented truth groups
jsonc
1
2
3
4
5
6
7
8
9
10
11
12
13
14
// Objective 20 goal: Phase 2 candidate pairs
// Suite: objective20_thorough, 95 shards
{
  // truth groups the candidates leave in 2+ pieces, must be 0
  "phase02_fragmented_truth_groups": {
    "baseline": 4,
    "bound": "== 0 on every shard and in aggregate"
  },
  // candidate rows / images, lower is cheaper
  "phase02_candidate_rows_per_image": {
    "baseline": 14.93,
    "bound": "below roughly 12.0"
  }
}

Phase 3: Edge scores

The pairs that mattered most were true bridges.
Missed true bridges fell from 442 to 109, and cross-truth rows fell from 3,219 to 1,973.
jsonc
1
2
3
4
5
6
7
8
9
10
11
12
13
14
// Objective 21 goal: Phase 3 edge scores
// Scorecard: core stratified40 + stratified80
{
  // pairs a group needs to stay whole, left unaccepted, lower
  "missed_true_bridge_edges": {
    "baseline": 442,
    "bound": "<= 110"
  },
  // pairs from different truth groups accepted, lower
  "cross_truth_above_threshold": {
    "baseline": 3219,
    "bound": "<= 2500"
  }
}

Phase 4-7: Grouping

The target was exact groups (ie. accuracy)
Accuracy rose from 97.88% to 99.4%
jsonc
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
// Objective 4 goal: Phase 4 initial grouping
// Scorecard: cached stratified40+80 core, 176,260 groups
{
  // exact groups / all truth groups, higher
  "score": { "baseline": 0.9788, "bound": ">= 0.9925" },
  // truth groups reproduced file for file, higher
  "exact": { "baseline": 172528, "bound": ">= 174939" },
  // predicted groups holding another group's files, lower
  "false_merges": { "baseline": 2438, "bound": "<= 2438" },
  // truth groups split across predicted groups, lower
  "false_splits": { "baseline": 1450, "bound": "< 1450" },
  // groups both split and merged, lower
  "mixed": { "baseline": 156, "bound": "<= 156" },
  // lowest score on any one shard, higher
  "worst_core_shard": { "baseline": 0.9742, "bound": ">= 0.99" }
}

Optimizing the Speed of the Algorithm

The challenge gave the whole run 60 minutes on 16 vCPUs with no GPU, so speed was part of the algorithm.

The Process

Once a phase worked correctly, I did not know how to make it fast, so I split the job between two agents.

I asked one agent how the phase could be optimized and collected its ideas.
I handed those ideas to a second agent to implement with the constraint the output of a given phase had to be byte identical

Using this approach, I was able to cut the runtime by 80-90% depending on the size of the dataset

Optimizing the Experiments

When stepping into the world of AI assisted research, the real constraint is the number of experiments you can run in a given time.

I needed to find ways to greatly speed the rate of experimentation while be constrained to online the compute on my single machine I had at the time.

Caching Every Phase

Phases 1 to 3 cached their output, keyed by the settings that produced it.

This allowed me to save massive amounts of time by not having to recompute previous steps, so agents could spend time experimenting on a give phase
Phase Caches and Time Savedsandbox

Moving Sweeps to the GPU

I don't know how to write GPU optimized code in general, let alone having it work on MLX... but agents do

The agents ported Phases 2 and 3 to MLX, Apple's machine-learning array framework, so sweeps could run on the one Apple GPU I had.

Shown below is a high level of the extremely complex code an agent wrote to run an optimize on my laptop
The agents wrote and ran hundreds of scripts.
The script is run_objective21_mlx_proxy_sweep.py, 1,518 lines.
The agents wrote it for Objective 21, to meet the Phase 3 bridge gate.
One run tested 50,761 configs against 6,493,245 pair rows on one Apple GPU.
That is about 330 billion checks in 2,733 seconds.
Call stack
main()Entry point. One run is one full proxy sweep
load_mlx(require_mlx=True)Imports mlx.core, requires Metal, makes the GPU the default
load_feature_arrays(feature_table)133 cached columns for 6,493,245 image-pair rows
load_configs(config_matrix, ...)50,761 rule configs become float32 threshold columns
cpu_baseline_by_shard(features, ...)Baseline counts per shard, used later for deltas
run_full_sweep(mx, features, ...)Configs outer loop, rows inner loop, sums counts per config
evaluate_chunk(mx, features, configs, ...)Called 650 times: 50 config chunks x 13 row chunks
make_feature_getters(mx, features, r0, r1)Each feature slice becomes a column (rows x 1) on the GPU
make_config_getters(mx, configs, c0, c1)Each threshold becomes a row (1 x configs), so they broadcast
accepted = (...) & ~global_brakeThe decision for every pair under every config at once
mx.matmul(shard_one_hot.T, mask)Per-shard counts as one matrix multiply
mx.eval(*aggregate_outputs.values(), ...)Runs the lazy MLX graph on the GPU and returns counts
result_rows(config_rows, ...)Missed bridges, cross rows, worst shard, rank score per config
pareto_rows(results)Keeps configs nothing beats on both missed and cross
write_json(..., benchmark_summary)Records device, shape, elapsed time, throughput, memory

Cleaning the Dataset

During the challenge, I only worked on the algorithm. I was disappointed with my result, so I kept going afterward and looked at the errors the model was making.

The errors pointed back at the data. The dataset was badly broken, so I had to fix it before the system could get any better.

I used AI for the scale and kept myself in the loop for the judgment calls.

Gemini, through the Batch API, tagged every image with its room, camera view, and quality problems.
I had agents build scans that removed duplicate images and near-blank, low-information frames.
A second Gemini pass looked at every group of images and flagged the ones that looked wrong.
I built a custom labeling studio to go through those flags, classify images, and correct groups.
Over the whole project, I spent roughly 8 to 10 hours labeling by hand, in batches.
My labels corrected the AI's classifications and fed back into each round of cleaning.

The cleaning ran in five stages, and the last two repeated until the dataset was clean.

Tagging Every Image

I started by having Gemini 3.1 Flash-Lite label every image, with BAML defining the label schema, through the gemini batch API

I was able to label all 250K images for around $100
typescript
1enum IndoorOutdoorTag {
2 INDOOR
3 OUTDOOR
4 UNKNOWN
5}
6
7enum CameraViewTag {
8 GROUND_LEVEL @description("Normal human-height interior or exterior property photo.")
9 AERIAL_DRONE @description("Drone or clearly aerial property view.")
10 // ... ELEVATED_VIEW, DETAIL_CLOSEUP, LOW_INFORMATION, OTHER, UNKNOWN
11}
12
13enum AreaTypeTag {
14 KITCHEN @description("Kitchen or kitchenette.")
15 BATHROOM @description("Bathroom, washroom, shower, tub, vanity, or toilet area.")
16 // ... 31 more, from BEDROOM and OPEN_PLAN to WAREHOUSE, POOL, and FACADE
17 UNKNOWN @description("Area cannot be determined from the image.")
18}
19
20type QualityFlag = "blurred" | "dark" | "overexposed" | "underexposed" | "rotated_or_orientation_issue" | "partial_or_occluded"
21
22class ImageClassification {
23 indoor_outdoor IndoorOutdoorTag
24 camera_view CameraViewTag
25 area_type AreaTypeTag
26 secondary_flags SecondaryFlagTag[]
27 quality_flags QualityFlag[]
28 tag_confidence "high" | "medium" | "low"
29 short_reason string
30 // ... each field also carries an @description in the source
31}
32
33function ClassifyDatasetImage(img: image) -> ImageClassification {
34 client Gemini31FlashLiteLow
35 prompt #"
36 {{ DatasetImageClassificationPrompt(img) }}
37 "#
38}
39template_string DatasetImageClassificationPrompt(img: image) #"
40 // ...
41 - Choose camera_view by how the image is shot, not by what area is shown.
42 - Use quality flags only for visual defects, not for normal real-estate lighting or ordinary HDR variation.
43 // ...
44 {{ ctx.output_format }}
45"#

Lock the label schema before the full run.

I changed the schema after the first pass, which meant a second full pass over all 266,376 images.
Being lazy cost me $100

Look at the Data

Using the Tags to Find Problems

The tags turned 266,376 images into something I could filter.

Gemini flagged 1,204 images as rotated or wrongly oriented.
It also flagged off-topic and sensitive images, including nude photos, which I removed and will not show.
With more than 200,000 images, I could not have found them by hand.

Remove Low Information and Duplicates

Some images could never be grouped right, so the easiest fix was to remove them.

Duplicate copies
The biggest cut removed 4,993 images, 3,075 of them duplicate copies.
Near-blank frames
208 frames, 145 too white and 63 too black, carried almost no information to group on.

Fixing Broken Groups and Duplicates

The rest of the problems were whole groups, not single images.

Duplicate groups
A hash and pixel scan found 3,818 duplicate pairs in 2,649 clusters.
An exact copy in two truth groups can never be scored right, because one file cannot sit in both.
Mixed groups
One truth group held three different rooms, with five brackets each.
Other groups held a frame that belonged in the group next to it.
Review cases
Each problem became a case with a proposed fix, 4,206 in all, and 1,916 needed my review.

Review Until It Was Clean

I had Codex build a review studio so I could make each call fast. The UI is clunky, and the video below does most of the explaining.

Grid overlays
Lines on every frame show a small camera shift or rotation against corners, windows, and rooflines.
Stacked, enlarged frames
I could blow frames up and compare them next to each other.
Review flags
Flags such as different room or small angle change went back to the agent with each call.
Drag-and-drop regrouping
I could move an image into the group it belonged in.
Hotkeys
Keys 1 to 5 each ran an action, so I could move through hundreds of cases in a sitting.

Each batch of calls went back to the agent for the next round of corrections.

Reviewing error cases in the correction studiostudio-review-walkthrough.mp4
A walkthrough of the review studio: each error case, the reasoning behind it, the grid overlay for comparing frames, and the buttons that accept or change a grouping.

Each accepted round became a new dataset version, and the baseline had to be rebuilt before the next experiment meant anything.

From Review Round to New Baseline
From Review Round to New Baseline

Did AI Make Me World Class?

I went in asking whether AI could make me world class at something I did not know how to do.

Yes, it did.

I did not know how the algorithm worked or how to do this at all. I was still able to run these systems and churn through things at a level better than people who were world class at this.

The Results:

The winner scored about 98.6 percent
My last entry placed seventh.
My final algorithm reached 99.4 percent accuracy on the cleaned dataset of about 260,000 images, much larger than the challenge set.

I started knowing almost nothing about vision models and finished with a SOTA algorithm.

Wild times we live in.

For the goal, context, workspace, and current-state files that kept these runs on track, read Objectives: A Lightweight Way to Keep Agents on Track for Long-Running Tasks.