Inside a 3,056-rated farming agent

I used coding agents to build a farming policy for Kaggriculture. On August 22, it reached a live rating of **3,056.31**. The final version I submitted scored **1,647.7**. Between those two results sits the part of the project I want to explain: how an automated research loop makes progress, and how its tests can stop telling you what you need to know.

The agents proposed strategies, wrote Python, ran games, inspected losses, and kept or rejected changes. Some repairs held up on the same recorded games. Others improved average cash while losing matches the previous policy had won. A better-looking experiment report could still leave me with a worse competition result.

This is a reconstruction from saved code, experiment reports, public replays, and submission receipts. “Day 0” means our earliest dated development report, August 10. The rating is the highest returned score among 66 archived submission histories we checked; it is not a claim of first place or a complete all-time record. History evidence ↗

From a crop loop to a competitive farm

Kaggriculture gives two players a 30-day season, with 24 decision steps per day. We buy land, hire workers, grow crops and tend animals. Both farms sell into one market: more output can depress the price. Winning means finishing with more banked cash after inputs, feed and wages. Travel consumes worker time that could otherwise service or deliver production.

By Generation 1, the baseline grew carrots with several workers. Generation 1 sent 21 experimenters after different hypotheses: crops, animals, labor, land, and delivery. The promoted melon strategy increased mean cash on the same development panel from 6,588 to 32,350. It won 52 of its 60 held-out tournament matches. Generation 2 combined activities into a portfolio allocator and reported 70 wins in a different 80-match tournament. These tournaments cannot be joined into one learning curve.

Aug 10Explore the economics

Single crops, species, workers and routes. A simple panel supplies rapid feedback.

Aug 22Verify the live peak

Gen68 reaches 3,056.31. Surviving code matches the uploaded file hash.

Sep 17–18Repair complete programmes

Production, feed, wages and sales are evaluated together on fixed replay cases.

Sep 30Audit the final package

Broader controls, responsive opponents, runtime and extracted-archive checks.

The original starter-panel mean cash rises from a 6,588-coin carrot baseline through 32,350, 107,200, 161,296 and 178,890 coins across the first four documented generations.

Early gains on the starter panel. These tests were separate from the later competition matchups.

Early gains included splitting feed pickups so the first carrier did not drain supplies needed by the others. Real opponent replays then exposed a larger weakness: starter-agent tests rewarded farms that struggled against aggressive land expansion and shared-market competition. A good crop had to remain profitable after both players sold it.

Archived Gen68 live rating declined after its August 22 peak while mean own cash stayed comparatively stable.

Gen68’s live rating after its August 22 peak. Daily cash stayed relatively steady as the opponents changed.

The autoresearch starting point

Our original program.md explicitly describes an adaptation of Karpathy’s autoresearch. His reference setup lets an agent edit training code, run a five-minute training budget, measure validation bits per byte, and retain or reject the experiment. A fixed evaluator makes the comparison meaningful. The instructions and experiment log are part of the research system.

We substituted a farming policy and game evaluations. The LLMs wrote and investigated Python; the submitted policy did not call an LLM during play. Initially, one evaluation contained 20 games against the starter and 20 against the current champion. A candidate needed no errors, improved head-to-head margin, and at least the champion’s starter-panel cash.

Design choiceKarpathy’s referenceOur game-agent adaptation
Editable objectTraining code in one fileInitially one policy file; later a frozen source tree
FeedbackValidation bits per byteWins, retained wins, own cash and opponent-relative margin
Evaluation boundaryFixed preparation and evaluatorFixed panels and engine, plus explicit opponent-model assumptions
MemoryExperiment instructions and results logLogs, generation reports, rejected hypotheses and saved replays

Karpathy’s agent instructions establish a baseline before changes and preserve failed results in the log. His 2019 training recipe also argues for inspecting data and verifying a simple evaluation pipeline before adding complexity. Both ideas mattered here: every economic intervention needed a checkable mechanism.

The research architecture

The research system had two jobs: explore ideas concurrently and make one defensible decision about what to keep. A coordinator read the rules, the current champion, and the accumulated learnings, then assigned a distinct hypothesis to each experimenter. Each experimenter worked in an isolated worktree and returned a candidate policy with its results.

Shared research memoryRules and instructions · current champion · results.tsv · LEARNINGS.md
CoordinatorRead the evidence, identify a mechanism, assign separate hypotheses
Experimenter AProduction and crop economics
Experimenter BWorkers and delivery routes
Experimenter CFeed, wages and market effects

Example hypothesis lanes; each produces separate code and a report.

Freeze and evaluatePin source and cases → benchmark queue → recorded controls and responsive games
Review the complete resultRetained wins · recoveries · both banks · paid inputs · runtime · archive checks
KeepPromote the checked candidate; verify any combination again
Reject or holdSave the failure or missing evidence for the next research brief
↺ Update the experiment log and learnings, then repeat
A locally qualified policy still needs a separate upload decision and live feedback.

The later research loop: parallel coding, serialized measurement, and evidence feeding the next generation.

The diagram combines the original coordinator/experimenter roles with the stricter evaluation process used in September. Early generations used different compute arrangements. By September 28, recorded benchmarks went through one shared queue, with at most three policy workers inside the active run. This kept competing experiments from changing timing-sensitive results simply by fighting for the same machine.

The coordinator rechecked promising claims and compared candidates with the champion. If the champion changed while another candidate was waiting, that candidate needed another comparison. Combining two successful branches also required a fresh test: workers, fertilizer, inventory and shared market prices made their gains interact.

What actually ran during a game

The coding agents belonged to the research loop. The submitted policy was Python and made no LLM calls during play. It received the current observation, chose actions, and returned commands to the game engine.

Current observationPublic farm state and available information
Python policyEvaluate production, labor, prices and routes
Commands → game engineApply actions and return the next observation
↺ Repeat for each decision step · no LLM in the game

The checked benchmark wrapper rejected missing cases, duplicate episode/seat pairs, incomplete games and candidate errors. Later reviews examined ordered bank changes and actual paid inputs. A proposed wage saving had to account for the productive work being removed; a profitable harvest needed a funded delivery route.

We kept three kinds of evidence separate:

Recorded

Opponent commands come from a saved game. Useful for matched repairs; limited when a real rival would change plans.

Responsive

Opponent code acts on the changed game state. Fresh worlds help test effects beyond the repair panel.

Live

Accepted uploads, hosted status and completed competition games. Local results do not establish a live rating.

What measurably improved

On one historical panel of the same 82 cases, the accepted Route848 baseline won 61 games. By Demand859V3, that became 79 wins and three losses. Improvements involved funded production and service, feed and fertilizer custody, competitive sale valuation, and returning output before the season ended. This was a repeatedly used development panel, so its improvement establishes repairs on those cases rather than unseen performance.

Wins on an identical 82-case recorded panel increase from 61 to 79 across tested source versions.

Wins on the same 82 recorded games rose from 61 to 79.

A separate September 30 panel tells a less flattering story. Versions 251, 336, 340, 361 and 397 each won 90 of the same 120 games, although their cash and margins differed. Version 340 earned more own cash than the final source on this panel; later choices also considered other difficult cases and runtime. There was no universal ordering of “best.”

Five exact versions all win 90 of 120 games, while own cash and margin change by version.

All five versions won 90 of the same 120 games. More cash did not always mean a larger winning margin.

What we learned from leaders

We inspected public farm states, labor, production and deliveries. A recent outcome-selected sample contained 29 own games and 25 leader views across 19 games. On engine day 6, leaders’ median farms had 22 strawberries, seven cows, three geese and nine paid hands; ours had 12, four, zero and seven. On day 10, our median cash was 15,589 versus their 3,869.

Selected public leader farms show more early strawberries, cows, geese and hired hands, alongside lower day-ten cash balances.

Public leader farms invested more in production and workers early in the season. These are observations from different games, not a controlled comparison.

The hypothesis was earlier capital deployment, followed by production that workers could actually service and deliver. A rival’s farm composition alone was insufficient: delivery timing, shared prices, purchased feed and escalating wages could reverse its economics. We needed to copy a mechanism and test it, rather than copy a board arrangement.

Explore the saved games

Choose a win or a loss, scrub through the season, and inspect the farms. The icons represent actual saved public states, with exact cash and worker positions. Tap a cell for its details. The viewer opens on day 6; slide back for the opening moves. The clock uses engine days 0–29. Recorded commands are requests, not proof of completed sales.

Replay a saved game

Saved public observations · no simulation

Our bank—
Rival bank—
Margin—
Loading…

Our farm

Rival farm

🌾 Wheat · 🥕 Carrot · 🍈 Melon · 🍅 Tomato · 🍓 Strawberry · 🐄 Cow · 🐑 Sheep · 🪿 Goose
F Farmer · 2 Hands at this cell · diagonal fill: locked land

Select a farm cell to inspect its saved state.

Our command entering this state

Requested from the prior observation; completed fills are not inferred.

—
Watch or download the four 37-second videos

The silent clips show all 720 saved observations at 24fps, plus opening and final holds. Cash is never interpolated. Captions mark checkpoints; fullscreen makes the grid easier to inspect.

RECORDED LOCAL TEST

Recovering a loss: −15 → +753 coins

A local loss becomes a 753-coin win: 67,293 versus 66,540. Watch the final-day carrot delivery.
RECORDED LOCAL TEST

A last-day route in a 241-coin win

A preserved local win, ending at 107,511 versus 107,270. The change improves the margin by 85 coins.
RECORDED LOCAL TEST

The fragile control: a 46-coin win

A fragile local win survives: 106,486 versus 106,440. Average gains elsewhere would not compensate for losing this game.
ACTUAL HOSTED MATCH

A completed game, a 1,226-coin loss

An actual hosted loss: 75,156 versus 76,382. The game completed; paid workers also performed productive work.

The clips are silent reconstructions from pinned replays, rather than original screen recordings. No policy or game engine was rerun. Only public farm states and our recorded commands are displayed. Case evidence · Video checks and source hashes. Gen68’s peak has a verified rating record, but no corresponding full replay was found in the saved records; its video is therefore absent.

Anomalies that changed our decisions

The narrow-panel mirage. A repaired historical-policy branch improved from two to six wins on 15 difficult cases, while losing both winning controls. The same exact source then won only 21/120, compared with its parent’s 90/120: five losses recovered, 74 prior wins lost. The small panel found behavior worth studying; it could not qualify a replacement.

A candidate wins six versus two on a selected 15-case panel, but only 21 versus 90 on the complete 120-case panel.

A targeted test looked promising; the broader panel revealed 74 lost wins.

The average concealed a loss. Candidate 390 improved mean margin over 191 recorded cases and recovered one loss, but lost two previous wins. In one, our cash rose by 86 coins while the rival’s rose by 358. Positive average income did not satisfy win preservation; the candidate was rejected.

A strategy could finish and still fail runtime. Candidate 388 recorded six wins and ten draws against a responsive opponent, but two own callbacks exceeded our local one-second criterion. Profiling identified repeated retirement-market calculations and route construction. Version 397 compared shortest routes before building commands and skipped empty market hours. Equivalent-function checks preceded full games. Its maximum own callback on a fresh 16-game panel was 0.786 seconds. The frozen peer still overran and remained explicitly unqualified; this was not a hosted latency guarantee.

Saved-callback measurements improve after route and market optimization; a separate completed episode's total elapsed time exceeds summed own callback time.

Callback runtime and total game duration measure different things. Neither chart alone explains the hosted loss.

Same cash did not mean the same path. An initial archive check matched 12 of 13 full paths. One 20-coin seed purchase shifted four callbacks; state later rejoined and both final banks matched. A predeclared repeat matched all 13. We retained both observations: the repeat did not erase the mismatch or prove universal determinism. Anomaly evidence ↗

The final result, without rounding away the losses

The final 397 source was accepted as submission 56722213 at 23:54 UTC on September 30. Its original saved hosted-status check remained unconfirmed at the research cutoff; the later official result is recorded in the closing section below. Local evaluation covered 133 recorded cases: 93 wins, 40 losses, retaining every parent win with zero new recoveries. Own cash improved in 41 cases and was equal in 92: 3,125 additional coins in total, or 23.50 per game. That was a small measured gain.

Final397 changes compared with361: 41 cases improve, 92 stay equal, zero regress, with no new win recoveries.

Final397 gained cash in 41 of 133 recorded cases, with no new wins.

The priority remained all first-40 wins, then all first-60 wins, then margins strictly above 30,000. One source had to meet each target; wins from different branches could not be added together.

Explore the final 60-game cohorts

Recorded opponent commands · exact source 397 · complete coverage, no imputed games

Margin > 30,000 Other win Loss · Select a game for cash and margin

View the complete static margin chartBoth complete chronological 60-game cohorts show remaining negative margins, with a marker after game40 and a strict30,000 target line.

The target was not achieved. Cohort A ended at 37/40 and 46/60; B at 36/40 and 44/60. Only 11/120 margins exceeded 30,000. The live peak of 3,056 belongs to the archived Gen68 source, not this final upload.

The research loop produced real repairs and useful negative results. It also showed where our process became slow: repeatedly optimizing a familiar development set could yield small cash improvements while leaving strategic losses intact. A better next cycle would reserve unseen opponents earlier, test complete economic programmes, and rank hypotheses by plausible recoverable deficit. A launched run, a higher average, and a successful upload each answer different questions.

What the tests could tell us

A win means finishing with more banked cash than the rival. A recovery means winning a case the parent policy did not win. Gaining cash in an already-won game is useful, but it is not another recovered loss. That distinction explains why the final version could gain 3,125 coins across the recorded panel without adding a single win.

The final checks included 133 recorded games, 16 fresh responsive games, two known runtime cases, and 26 comparisons of the extracted submission archive. Those archive comparisons repeated existing cases; they were not 26 new opponents. The responsive panel used one frozen opponent, so it still offered limited evidence about how the policy would fare against the field.

I pinned the tested source, baseline, case list and submission archive so the results could be traced back to the code that produced them. The detailed evaluation notes, figure data, and failure analysis are available for readers who want to inspect the mechanics.

Ranking and closing results

The final upload completed, but it did not reproduce the historical 3,000-plus peak. A one-time official check at 8:14 p.m. Pacific on September 30 returned the results below. Submission scores, team placement and the archived peak describe different records. Download the dated closing evidence.

RecordVerified resultWhat it establishes
Archived Gen68 peak3,056.31August 22 live rating; no contemporaneous rank located.
Historical team placement13th · 2,943.4September 13, 6:02 p.m. Pacific snapshot; not the final competition placement.
Earlier Margin361 upload
56719676
2,145.4 · COMPLETEOfficial submission score at the closing check.
Final397 upload
56722213
1,647.7 · COMPLETEOfficial submission score at the same check; below the earlier upload.
Settled final competition rankUnverifiedOur team was absent from the returned first 200 leaderboard rows. That response does not establish an exact or settled final placement.
Final397 recorded cohortsA: 37/40 · 46/60
B: 36/40 · 44/60
Complete local panels; zero-loss targets remained unmet.

The 1,647.7 result is part of the outcome, even though the final local candidate preserved wins and gained cash. These two uploads did not play a matched live schedule, so their score difference cannot isolate the effect of the code change. The practical lesson is to judge promotion using broader responsive opponents and live feedback, while retaining the exact failures that the development panel missed.

References

  1. Andrej Karpathy, autoresearch, original README. The fixed evaluation instrument, bounded experiments and human-authored research programme informed our adaptation.
  2. Andrej Karpathy, autoresearch agent programme. Baseline measurement, experiment logging, and keep/discard decisions. Observed repository and file revisions are saved in reference notes.
  3. Andrej Karpathy, A Recipe for Training Neural Networks, April 25, 2019. Data inspection, simple evaluation baselines and incremental hypothesis testing. This is a methodology reference, not a separate autoresearch blog post.
  4. Kaggle, Kaggriculture leaderboard. The official closing submission/leaderboard API snapshot is bundled in dated results; final placement remains unverified.
  5. Project records: hashed documentation inventory, pipeline design audit, and development and live-rating evidence. These distinguish documented starter experiments, matched recorded panels and archived live history.
  6. Experiment analysis: exact figure datasets and anomaly evidence. The figure datasets cover all nine Matplotlib charts.
  7. Game visualizations: interactive replay provenance and video provenance. Both reconstruct saved public states; they are not new simulated results or original screen recordings. The video manifests also contain frame mappings and source hashes.

Results use records through September 30, 2026, including the official closing check at 8:14 p.m. Pacific. Supporting records are linked above.

Cite this post

If you reference this post in your work, please cite it as:

@article{nagabhushanaradhya2026kaggricultureautoresearch,
  author  = {Nagabhushanaradhya, Subramanya},
  title   = {Inside a 3,056-rated farming agent},
  journal = {subramanya.ai},
  year    = {2026},
  month   = {September},
  url     = {https://subramanya.ai/2026/09/30/kaggriculture-autoresearch/}
}
×