Flagship demonstration
We are Summit Ridge Roofing in Denver and we want more qualified leads. Right now we get about 40 leads a month and only half are real jobs. Get us to 60 qualified leads a month within 90 days, without spending more than $6,000 a month on marketing. Don't do anything that could hurt our reputation or our reviews.
A chatbot would have written a marketing plan. Nexefiy ran one — under the owner’s authority, against the owner’s rules — and learned from how it went.
Every step below went through the real API, runtime, database and row-level security. Each card lists what was persisted and what the step proves; the integration suite asserts every one of those claims against the database.
Summit Ridge Roofing has an organization, a ratified Constitution, 3 active policies, 7 registered capabilities and 18 observations from 6 of its own systems. No objective exists yet.
tenants#1governance_constitutions#1Read one sentence from the owner and extracted a lead generation objective: 40 leads/month today at 50% qualified → 60 qualified leads/month within 90 days, capped at $6,000/month.
Persisted the objective as a structured genome — 2 measurable outcomes, 1 constraint, 5 guardrails, 5 forbidden actions, 2 approval gates, 2 open uncertainties — and the owner activated it.
objectives#118 live facts about 6 entities. 1 contradiction kept open (both sources preserved), 4 stale facts detected. Overall confidence 0.772, critical coverage 100%.
Weighed 18 pieces of evidence from the World Model (13 verified), 7 information gaps and 4 load-bearing assumptions. Verdict: tentative (uncertainty 0.74) — not enough to conclude, so the plan must buy information before betting on it.
uncertainty_assessments#1Three beliefs the plan depends on became falsifiable hypotheses, each with a real (not simulated) experiment and a decision threshold. Beliefs start at their evidence, not at optimism.
hypotheses#1hypotheses#2hypotheses#3Generated 5 distinct strategies from 7 levers — including the do-nothing counterfactual and the tempting shortcut, so each can be judged on the record rather than silently dropped.
The compiler chose a hybrid architecture (complexity medium, confidence 0.897): A medium-complexity objective spanning agents and external systems → 4 participants. It formed 4 entities and an objective graph of 8 nodes with a governance gate. Re-compiling the same input returned the same stored record.
compilations#1intelligence_entities#1intelligence_entities#2intelligence_entities#3intelligence_entities#4Granted 6 bounded authorities — each names one capability, one implementation, a risk ceiling and a budget cap; medium-risk ones also require a human approval. The authority pipeline then allowed Google LSA for the Lead Volume Agent and denied the list vendor at the authority stage — no entity holds a grant for it.
capability_authority_grants#1capability_authority_grants#2capability_authority_grants#3capability_authority_grants#4capability_authority_grants#5capability_authority_grants#6Judged 11 distinct actions against the Constitution and policies: 2 allowed autonomously, 8 need a named human's approval, 1 blocked. 1 strategy removed before it could be simulated or run.
governance_evaluations#1governance_evaluations#2governance_evaluations#3governance_evaluations#4governance_evaluations#5governance_evaluations#6governance_evaluations#7governance_evaluations#8governance_evaluations#9governance_evaluations#10governance_evaluations#11Simulated 4 admissible strategies (500 seeded draws each, spread widened by how unsure each belief is) and ranked them on the pessimistic case. Chosen: “Quality-first mix with a learning budget”. The cheapest-leads plan has a higher average (86.23) but a p10 of 23.34 — it rests on a stale, low-confidence estimate, so it lost.
simulation_runs#1simulation_runs#2simulation_runs#3simulation_runs#4simulation_comparisons#1Started one durable run (a replayed request returned the same run). 6 tasks executed through the admission guard: 4 waited for and received the owner's approval, the others ran on standing authority. The ads API rate-limited 1 attempt; the runtime retried per the task's policy. Spend $5,919 of a $6,000 cap.
execution_runs#1run_tasks#1run_tasks#2run_tasks#3run_tasks#4run_tasks#5run_tasks#6Pulled the results from the run's persisted task outputs: 82 leads, 51 qualified. Wrote each channel's measured qualification rate into the World Model (2 contradictions with prior belief detected and resolved to the measurement), and recorded the actual next to the simulation's prediction (74.26 predicted → 51 actual).
world_facts#1world_facts#2world_facts#3world_facts#4simulation_actuals#1Pilot month: 51 qualified leads against a target of 60 (78% of the way from the baseline), $116 per qualified lead against $100. Not yet met — and the experiments explain why: the three hypotheses were judged by observed measurements.
evaluations#1Stored 11 memories from the pilot: the episode itself, 6 confirmed beliefs, 3 pieces of negative knowledge, and the playbook. Retrieval for “what does not work” returns the negative knowledge first — so no future plan repeats the mistake.
memory_entries#1memory_entries#2memory_entries#3memory_entries#4memory_entries#5memory_entries#6memory_entries#7memory_entries#8memory_entries#9memory_entries#10memory_entries#11Negative memory removed meta_lead_forms from the next portfolio (5 options); corrected beliefs reallocated the budget to storm_landing_pages $3,300 and google_lsa $2,700. The Lead Volume Agent's change went through the Evolution Lab — sandbox, simulation, evaluation, the owner's approval, a reversible deployment — and month two delivered 66 qualified leads at $90 each: target met, within the cost target.
simulation_runs#5evolution_experiments#1intelligence_entity_versions#1execution_runs#2evaluations#2tests/integration/flagship runs the whole loop against PostgreSQL and PostgREST — the same engine Supabase runs — through the real middleware and route handlers. CI replays it on every change and fails if the result differs from this page by a single character. Regenerate with npm run demo:roofing.
No real ads were bought. The effects of executed actions — leads from Google, Meta, landing pages and referrals — come from a deterministic market sandbox behind the runtime’s tool seam, whose true numbers Nexefiy never sees while planning. Everything else — the objective, world model, uncertainty, hypotheses, compilation, grants, governance, simulations, approvals, observations, evaluations, memory and evolution — is the production system.