FutureMaster · Technical edition

For kids who already ship: the 12-week operating spec.

## Who this version is for

You already have agency. You've deployed something, or you're one weekend away from it. You know what a repo is, you've argued with a model about its own hallucination, and you don't need to be sold on AI mattering.

What you probably don't have: a systematic practice. Most technically-inclined kids plateau at "pretty good at prompting" because nobody hands them the actual discipline: mission design, context architecture, verification protocols, model evaluation methodology, and the deployment cadence that compounds. This document is the spec for that discipline.

// Most technical kids plateau at pretty good at prompting. Nobody handed them the discipline.

## The core claim, stated precisely

Model capability is commoditizing. Marginal cost of output (code, prose, media, analysis) trends toward zero. Under that condition, the residual sources of human value are:

// An agent is not a model. It is a mix: model + memory + tools + orchestration.

> 1. Problem selection

deciding what is worth building. No objective function exists for this; it requires domain obsession.

> 2. Specification quality

the fidelity with which you can transmit intent to an execution system. This is an engineering skill, not a writing skill.

> 3. Verification

detecting when a fluent output is wrong. Asymmetric: generation is cheap, validation is scarce.

> 4. Taste

discriminating between competent output and exceptional output, and knowing how to force the gap closed.

> 5. Distribution

moving other humans. The one market AI cannot flood, because trust is human-denominated.

The 12 weeks are structured training on exactly these five, in that order. Everything else (syntax, tools-of-the-month, prompt tricks) is deliberately excluded as depreciating assets.

## Architecture of the FutureMaster program

$ Flywheel: Spark -> Tools -> Ship -> Share. You run the full loop at least twice. The loop is the unit, not the week.

$ Cadence contract: weekly expert session (practitioner, not lecturer, ends with a challenge), 2x build blocks, small-group demo day, minimum one closed-laptop explanation, and a teach-down to the next age group. Teaching down is not charity; it is the cheapest known test of whether you actually compressed the concept.

$ JARVIS: the advanced AI agent built into FutureMaster, a persistent instance bound to you for the duration. Architecture is inspectable by design: model + memory + tools + orchestration. The engine rotates under it (you will watch this happen); the memory and context persist. Two things are being trained simultaneously: you on directing agents, and your agent on you. Tuning it: context curation, memory hygiene, behavioral configuration: is an explicit skill track. You are not using an assistant. You are operating and training one.

$ Dynamic curriculum: JARVIS holds the program arc plus your project state (spark, capstone edges, debrief history) and re-renders each week's challenge into your lane. Same 12-week invariant for everyone; per-kid execution path.

## The ladder (progression model)

Passenger -> Typist -> Director -> Owner -> Leader.

You likely enter at Typist, possibly early Director. The program's claim: most self-taught technical kids are weaker than they think at the Director rung because they've never been graded on brief quality, only on eventual output. Weeks 3-6 will locate you precisely.

## Phase-by-phase, with the actual mechanics

## Weeks 1-2: Spark (don't skip this because it sounds soft)

The spark phase looks like the fluffy part. It is the load-bearing part. Problem selection (the #1 residual skill above) is impossible without a domain you involuntarily think about. The onboarding extracts it, the Spark Statement formalizes it (your lane, why it pulls, three things you want to exist that don't), and JARVIS ingests it as the seed context for everything after.

Week 2 formalizes the director's loop: brief -> review -> reject -> own. Your mission briefs run through the Prompt Review analyzer and get graded structurally. The asking-vs-directing distinction is made mechanical: a command pins the model to your context, constraints, and definition of done; an ask invites it to improvise from its priors, which is where hallucination lives.

Baseline capture: closed-laptop explanation of your most recent project, recorded. This is the "before" leg of the program's primary before/after diff.

## Weeks 3-6: Tools (the dense phase)

// The unglamorous 10x. Deliverable: a reusable context briefcase for your lane (goals, constraints, taste references, prior work), plus an A/B: same task with and without the briefcase, difference measured. Numbers, not vibes. Context management across long-running work: gathering, structuring, preserving: is treated as a first-class engineering discipline, because it is one and nobody teaches it.

## Week 3 - Context engineering

// Live demo: a coach induces a confident, fluent, wrong output in your domain. Then the drill: source, seam, sanity-check. You build a wrong-log: every hallucination caught, documented against primary sources: and wrong-log depth becomes a tracked stat for the rest of the program. Trust architecture gets explicit: always-verify (numbers, citations, anything you sign), sample-verify, never-delegate (the final call). The transferable asset: you stop reading confident output as fact, permanently, across every information surface you touch.

## Week 4 - Verification protocol

// Mission design as a spec discipline. The ten laws, which you'll recognize as good systems engineering applied to agent delegation:

## Week 5 - The One-Shot Canon

$ 1. State the mission, not the steps (declare the goal state, let the system search).

$ 2. Grant autonomy explicitly (define its decision boundary or it blocks on you forever).

$ 3. Define done as evidence (a link, a passing test, a number: not "looks good").

$ 4. Front-load context handles, not dumps (pointers over payloads).

$ 5. Declare priorities for trade-offs (fast/cheap/perfect: rank before it guesses).

$ 6. Set budgets and stop rules (time, money, attempts; name the abort signal).

$ 7. Plan before payload (strongest model distills spark -> plan hierarchy: north star, phases, tasks, checks: before any build starts).

$ 8. Fence the invariants (name what must not change, or it will change it).

$ 9. Batch the whole ask (one complete brief beats twelve clarifying rounds).

$ 10. Debrief into memory (bank the lesson; pay the toll once).

Tracked stat: one-shot success rate: shipped with zero follow-up corrections. Climbing the size axis (paragraph -> page -> prototype -> full mission) at constant one-shot rate is the game.

// Domain scan, problem mapping, human interviews, opportunity sizing, pick ONE. Six lanes: Build, Creative, Research, Venture, Community, Advocacy. The pitch is graded harder than the eventual solution, deliberately: the choice is the skill. No capstone starts until the mission lands with the room. Hard gate.

## Week 6 - Problem Selection Studio (midpoint gate)

## Weeks 7-10: Ship

// Ship Zero: smallest real version of the capstone, public by demo day. Deployment is defined broadly but strictly: URL, published piece, listed product, sent pitch: it must touch reality. Ship cadence tracking starts and runs to week 12. Feedback does not exist pre-deployment; everything before is simulation.

## Week 7 - Deploy

// Execution is commoditized; discrimination is not. Structured critique reps: study the acknowledged greats of your lane, then push AI output to the standard they set in you. The 10x Pass: take Ship Zero, make it 10x better and unmistakably yours, and log every decision the machine could not have made. The decision log is the deliverable as much as the improved artifact.

## Week 8 - The taste pass

// The mix model made concrete: model + memory + tools + orchestration, each component's failure modes examined. Then the Model Race: identical task, two frontier models, same agent scaffold. Compare on output quality, speed, factuality, taste, failure modes, tool-use behavior, and human-correction load. You keep a Model Scorecard of your own evals. The meta-skill: when the next frontier model drops, you can characterize it in days: strengths, weaknesses, personality, where it slots in your stack: instead of consuming takes. Hype immunity through personal eval infrastructure.

## Week 9 - Agent systems

// Cold brief, unseen problem, 20 minutes, live, paired (one performs, one observes with a rubric, swap). Sample briefs: load-calc a bridge concept to destructive test; re-route hurricane logistics around a closed bridge; find three growth experiments for a flatlined 96-WAU browser extension; cut 60 seconds of teaser from an oversized lore bible; design the survey and argument that moves a school principal; turn a WWII supply-line dataset into an interactive map for non-experts. Everything from weeks 2-9 under time pressure: brief quality, verification under load, recovery from bad output, honest ownership of the result.

## Week 10 - The Director's Exam

## Weeks 11-12: Share (the part technical kids skip, at maximum cost)

The argument, stated without softness: when everyone has access to the same superintelligence, the machine side of every field flattens. The residual competition is entirely human-side: trust, teaching, recruiting, moving a room. Technical kids who skip this cap their own ceiling regardless of build skill.

// Cold Ten: ten real outreach messages to professionals in your lane (mentorship, feedback, or customers), response rate analyzed like a campaign. Plus the Human Edge thesis: pick a great you admire, name three strengths, run AI against them honestly, and write up the capability map. Concluding that AI closes the gap somewhere is a legitimate result; the skill being graded is honest capability mapping, not a predetermined answer. Plus one teach-down.

## Week 11

// Capstone ships in final form. Portfolio assembled: 3+ live creations, wrong-log, mission briefs, decision logs, outreach record. The Closed-Laptop Moment: live presentation to parents, mentors, and external practitioners: the problem, why it mattered, how AI was leveraged, where the human-only value went in. Laptop closed for Q&A. The "after" recording is diffed against week 1's "before."

## Week 12

## Measurement (the actual grading function)

Output is not graded: AI made output free. The tracked variables are the scarce ones:

$ Closed-laptop before/after: ownership; anti-slop

$ Ship cadence W7-12: bias to real

$ One-shot success rate: spec quality

$ Wrong-log depth: verification reflex

$ Pitch score: problem selection

$ Capstone (external judges): the whole stack

$ Humans moved: distribution skill

$ Ladder rung: overall progression

Every claim in your work needs provenance a human checked. Evidence over vibes is not a slogan here; it is the assessment mechanism.

## What is deliberately excluded

$ Prompt tricks. They expire with every model release. Invariants (judgment, briefing, verification, taste, agency) compound instead.

$ Coding syntax. Contrarian but reasoned: you have a fleet of world-class software engineers on tap awaiting a clear mission. The scarce skill is decomposition, delegation, specification, and system connection: what actually made the great coders great. If you already code, nothing stops you; it is simply not what is being trained.

$ Tool-of-the-month. Skills anchor to components and systems, not product names. The engine under JARVIS rotates on purpose.

$ Output grading, abstract ethics, theory without shipping. Ethics appears with teeth: real cases, argued from multiple positions, plus a personal strategy memo: where you will be irreplaceable in 2035, and why.

## The compounding argument (why 12 weeks and not a YouTube playlist)

Each loop deposits three non-transferable assets: a live creation, a sharper judgment, and the identity of someone who finishes. Twelve weeks bootstraps the loop under coaching pressure with a community where shipping is the norm. A decade of the loop produces something historically new: one person with the agency to do what previously took a team, a budget, and a decade of permission. A human who directs the machines, moves other humans, and never stopped caring about the thing.

That is the spec. The rest is reps.

That is the spec. The rest is reps.

// applying = joining the vetting queue. we reply to fits.

FutureMaster · model + memory + tools + orchestration

The Operating Spec | FutureMaster