Every agent has a job, an income, beliefs, and a memory of what you showed them last time. Show them a policy and read how a country reacts. Show them a product at a price and read what a market will pay. You get a decision memo or a market research report — not a guess.
Every agent answers a decision in its own words, and the reasoning is the point — it’s what a percentage can never give you. A sample of what the population said about one decision:
The engine underneath is the same: people with lives, reacting to something specific and saying why. What changes is what you hand them, and what comes back.
A policy, a memo, a press release — anything with consequences. Every agent reads it and weighs it against their own life before answering.
The same population becomes a consumer panel. It reads what you are selling and what you want to charge, then answers whether it would buy.
Whatever you put in front of the population, it answers as itself and says why. These are the rooms teams walk it into — a cabinet, a war room, a pricing meeting, a launch review — each ending in something you can hand to someone.
Pre-test a reform, subsidy, or regulation on a country before it is announced. Read the coalitions that form and the objection you will have to answer first.
Rehearse a hostile news cycle. Attack angles, a coalition-risk map, and a language-distortability scan show how a statement gets reframed — before it goes out.
Find the number a market will actually pay. Van Westendorp, a demand curve, elasticity, and a revenue-optimal price — computed from the panel, not guessed.
Concept and ad testing, segmentation, positioning, ranked channels, and a tiered influencer plan — packaged as a launch-ready market research report.
Model how a restructuring or a return-to-office mandate lands across departments, and where the resistance concentrates, before you send the all-hands.
Run competing framings of the same decision past group representatives and watch opinion move round by round. Keep the wording that actually shifts people.
A cabinet weighing a reform, a founder before a build, a fund in diligence, an agency before a shoot — each walks up with a different question and the same fear of finding out too late. The proof line under each card is the published research, not our own scoreboard.
Drop a policy, subsidy, or regulation on a simulated country and read the coalitions that form and the objection you will have to answer first — while it can still be reworded.
Fine-tuned agents predicted individual policy preferences at up to 77% vs ~51% random — Royal Society, 2024
Screen concepts, probe pricing, and stress a product-market-fit claim against 10,000 panelists with incomes and current alternatives. The report is allowed to come back no-go.
Synthetic interviews run at ~$2–60 each vs $80–120 per agency-recruited human respondent
Test messages and creative across segments and markets, screen five concepts down to two, and keep only the framing that moves opinion — before a media budget is committed.
Replaces $25k–$100k, 2–4-week creative pretests — so you test 10 variants, not 2
Re-run a founder’s demand and pricing claims against the segment they are targeting. You get the questions to ask in diligence — not a verdict, and never a substitute for the round.
EY reproduced a six-month global study in one day at 90% median correlation across 53 questions
Pre-test storyboards and scripts against the audience they are meant for, then rehearse the social reaction and the news cycle in the war room to see how a line gets reframed.
One PR firm tested six candidate narratives with 189,756 synthetic responses before going public
Run a script, trailer cut, or episode past audience segments and compare versions on reaction and reasoning. Built for choosing between cuts — not for predicting box office.
A synthetic focus group matched live family answers >95% of the time in one published study
Building the country and generating the people happens once, whichever question you are asking. After that the pipeline runs on its own. Where a phase relies on a published model, it’s named — the coefficients are inspectable, not hidden behind a prompt.
Paste articles, reports, or a handful of facts. Country Studio returns a structured dossier: demographics, sectors, media, fault lines.
50 to 10,000 agents — each with a name, profession, income, beliefs, and a memory that carries between runs.
A policy, a memo, a press release. Every agent reads it and weighs it against their life.
Each agent weighs personal impact, exposure, identity, and institutional trust, then lands on support, oppose, neutral, or conditional.
Agents with shared interests cluster into stakeholder groups. Business owners, religious leaders, and local influencers surface on their own.
Representatives persuade; neighbours argue. Echo chambers and consensus emerge from the models rather than being scripted.
Coalitions, fault lines, the sectors that absorb the shock, and the objection you will hear first.
A live web sweep compiles the dossier: price bands, named competitors, perception themes, and a category map. Every claim carries the URL it came from.
A strategy brief states the product-market-fit claim, the draft ICPs, the price hypotheses, the positioning territories, and the risks — before anyone reacts. Written afterwards, they could only be rationalised.
Up to 10,000 consumers answer at a given price: would they buy, would they switch, what is their ceiling, and what do they use today.
Van Westendorp, demand, elasticity, and the revenue-optimal price are computed, not written — no model does the arithmetic. k-means then clusters the panel on its own answers.
Twelve sections, a verdict on every hypothesis the brief pre-registered, and a self-contained HTML file with the charts inlined.
Paste articles, write a handful of facts in your own words, or drop in an image. The studio synthesizes a structured dossier, then injects it into every agent-generation, media, and propaganda prompt downstream. A richer dossier is a more lifelike simulation — and the same tool builds a real country or one you invent.
The population blueprint every agent is drawn from — age, income, profession, and the economic sectors that absorb a shock.
The real publications and their leanings, so simulated coverage reads like the actual press rather than a generic newsroom.
The tensions, taboos, and specifics of the place — the context an agent needs to react like a local, not an average.
Who opposes whom and who is allied with whom — the entity graph that seeds lobbies and powers the risk scans.
This is the deliverable a launch decision is actually made from. Sizing figures are marked estimated rather than measured, and their assumptions are printed next to them — that is the arguable part, which makes it the part worth reading.
Go, conditional-go, pivot, or no-go — with the confidence level stated and the reason it is only that high.
TAM, SAM, and SOM, labelled estimated rather than measured, with the assumptions boxed so you can argue with them.
Named competitors with prices, share estimates, strengths, weaknesses, and a threat level each.
Price against perceived premium. A quadrant with nobody in it is either a gap or a warning; the white-space note says which.
Van Westendorp. Where the four curves cross bounds the range the panel finds credible.
Share of the panel whose ceiling clears each price, with the prices you actually ran marked as stronger evidence than the interpolated line.
k-means over the panel’s own answers, drawn as a force-directed map. A model names the clusters it did not choose.
A recommended territory, an archetype, a promise, reasons to believe, messaging pillars, and what to avoid.
Ranked lead, support, and test channels, each with a rationale and the segments it reaches.
Nano, micro, and mid tiers with follower ranges, creator archetypes, content angles, platform mix, budget split, and guardrails.
Ranked reasons to buy and reasons not to, each carrying verbatims from named panelists.
Validated, rejected, or inconclusive on every claim the brief pre-registered, with the evidence and the so-what. Then a phased go-to-market and the questions this study could not settle.
The whole report exports as one self-contained HTML file with the charts inlined — it opens anywhere and prints to PDF. The panel responses, the segments, and the competitor set come out as CSVs.
A real study of a meetup app came back no-go, with four of its five pre-registered hypotheses rejected. It also found that price was never the problem. Because the claims were written before the panel ran, a rejection is a finding rather than a mistake — the study was set up to be able to lose.
Every agent carries a name, a profession, an income, and a memory that carries between runs. Facing a policy they hold a stance. Facing a product the same person is a panelist with a budget, a routine, and something they already use instead. A sample from the current run:
Each node is one agent with its own name, profession, beliefs, and memory. In a policy run the lines are who reacts to whom when a decision lands; in a market study the same people are the panel, and a second map groups them by how they answered. Four thousand are shown here; the engine scales well past that. Drag to rotate.
Every number in either report traces back to one of these. The first four decide how a country moves; the last four decide what a price is worth. All eight are arithmetic — a language model names things it did not compute. If you disagree with a coefficient, you can change it and run it again.
P(k) = exp(Uₖ) / Σⱼ exp(Uⱼ)
How each agent picks support, oppose, neutral, or conditional from calibrated utilities.
xᵢ(t+1) = λᵢ·xᵢ(0) + (1−λᵢ)·Σ wᵢⱼ·xⱼ(t)
Each agent blends stubbornness with neighbour influence until the population settles.
xᵢ(t+1) = mean{ xⱼ : |xᵢ−xⱼ| < εᵢ }Agents only listen within their confidence bound — echo chambers emerge on their own.
Δx = (I − A)⁻¹ · Δd
A shock in one sector cascades through the ten-sector economy via the inverse matrix.
OPP = { p : F_cheap(p) = F_expensive(p) }Four cumulative curves over the panel’s price ceilings. Where they intersect bounds the range the market finds credible.
ε = (ΔQ / Q) ÷ (ΔP / P)
Read off the demand curve at the tested price. Below −1 the market is price-sensitive; above it, the price is not what is stopping them.
p* = argmax₍ₚ₎ p · D(p)
Demand at each price times that price. Deterministic arithmetic over the panel — no language model computes it.
argmin₍S₎ Σₖ Σ₍x∈Sₖ₎ ‖x − μₖ‖²
Clusters the panel on what it actually answered rather than on demographics you picked in advance.
The idea that people can be simulated well enough to decide from is a research finding, not a pitch. These are the peer-reviewed results the method rests on — figures from the field, cited so you can check them, not our own marketing numbers.
Agents built from two-hour interviews with a stratified sample of 1,052 Americans answered held-out survey questions at 83–86% of each person’s own two-week test–retest consistency — the practical ceiling. Thin demographic personas managed only 74%.
Conditioning a model on real socio-demographic backstories reproduces the response distributions of distinct human subgroups — the founding result that a population, not just an average, can be simulated.
Given endowments and preferences, LLM agents qualitatively replicate five classic behavioral-economics experiments — fairness, status-quo bias, social preferences — the basis for reading demand-side behavior off simulated people.
Eliciting free-text reactions and mapping them to a rating scale — rather than asking for a number outright — reaches about 90% of human test–retest reliability, validated across 57 real product surveys and 9,300 respondents.
Augmenting a small 5% human sample with LLM-simulated preferences lifted aggregate estimation from about R²≈30% to ≈75% — the hybrid that makes simulation a supplement to real fieldwork rather than a replacement for it.
Put a policy in front of a country and read the memo. Put a product at a price in front of a market and read the report. Either way you end up holding a document you can hand to someone.