AI will kill us all! This is what we think when we hear about yet another AI breach by AI swarms. I know that fear is selling, but still, we have to step back and understand what actually happens.
AI is a powerful tool, but it’s still just a tool. In this article, I’d like to explain a little bit more about what AI swarms are and to demonstrate that you can do bad and good things with them, like every tool in human history.
So, first things first. What is an “AI Swarm”?
An AI swarm is a group of AI agents that work together in a coordinated way to solve complex tasks, like social insects such as ants or bees. This approach leverages collective intelligence to achieve better outcomes than individual agents could accomplish alone.
The term “swarm intelligence” was first introduced in 1989 by Gerardo Beni and Jing Wang in research on cellular robotic systems. It described how simple, decentralized agents could cooperate to produce intelligent group behavior.
Over time, more and more researchers approached the field, and with the invention of large language models, AI swarm intelligence became popular.
I decided to put aside all doomsday theories and think of a way to exploit the capabilities of swarm intelligence in real business scenarios.
This is my story of using many agents to audit my sales page - a task that would usually require a lot of visitors and a solid budget.
The AI swarm intelligence for business
Every sales guru tells you to “validate your positioning before you scale.” Nobody tells you how to do that when you have no audience yet, no budget for a research panel, and a sales page that’s been live for three weeks with numbers too small to mean anything.
So I tried something else. I built a simulation that runs locally, fed it my own market research and my own sales copy, had it build a crowd of AI characters out of that material, and let them argue about whether they’d buy from me.
The short version: it told me my own customer profile was wrong. Not slightly wrong — wrong about the thing that actually drives the buying decision. That cost me a weekend and about 1,200 API calls to find out, which is a good deal cheaper than finding it out from an ad budget.
This is a field report. It covers what the setup actually involves, what the agents got right, what they confidently made up, and whether I’d do it again.
What the thing actually is
The engine is called MiroFish, an open-source multi-agent simulation project that builds digital worlds to predict the future. You basically upload seed materials and describe you prediction requirements in natural language. The AI prediction engine will return a detailed report and deeeply interactive high-fidelity digital world.
MiroFish was created by Guo Hangjiang (known as 666ghj on GitHub), a senior undergraduate student in China. The project has gained significant traction and is backed by strategic support and incubation from Shanda Group, a Chinese investment firm. Under the hood, MiroFish’s core simulation architecture is powered by OASIS (Open Agent Social Interaction Simulations), an open-source engine developed by the CAMEL-AI research community.
You can use the Mirofish AI prediction tool as a paid service, or you can host it yourself.
I run the English version called MiroFish-Offline in a virtual machine on my own hardware rather than as a hosted service, for reasons I’ll get to.
The pipeline is five steps:
- Work out the vocabulary — the AI reads your documents and works out what kinds of things they talk about (people, problems, products) and how those connect.
- Build the map — it breaks the documents into pieces and builds a map of what connects to what. It creates a graph database out of these documents. Neo4j is the name of the graph knowledge database that holds the map.
- Environment setup — the AI writes the characters from that map, rather than from a paragraph you typed. This is the part that matters.
- Simulation — the characters post and reply to each other over a set number of rounds.
- Report — a final agent analyses the whole run and predicts an outcome.
Step 3 is the whole reason this beats asking ChatGPT to roleplay a customer. The characters come out of a graph knowledge database, so they inherit its shape — including the parts you didn’t consciously notice you’d written down.
Architecture
Everything runs on a virtual machine except the reasoning. I used the Ollama API service for the reasoning due to my local hardware limitations:
| Component | Where it runs | Does your text leave the machine? |
|---|---|---|
Turning text into numbers the machine can compare (nomic-embed-text) |
Local Ollama container | No |
| The map of your documents (Neo4j) | Local container | No |
| The AI reasoning steps (reading, extracting, writing characters, the report) | Ollama Cloud | Yes |
This is worth being explicit about, because it’s the kind of detail that usually gets skipped. Keeping that step local protects the numerical index, not the actual words. Every reasoning step sends text from your documents to a third party - even if Ollama claims that no logs are beeing recorded, the only way to be certain is an offline mode without any messages leaving your local setup.
My seed documents contain my pricing, my positioning and my ICP definitions — so this is a real trade-off, not a theoretical one.
I accepted it because my laptop can’t run a model strong enough to make the simulation worth doing. If your seed documents are client data rather than your own strategy, run the LLM completely offline.
The constraint that shapes everything
Keep the seed set to about three documents. It isn’t a hard cap, but past three the extras stop attaching to the graph reliably, and you end up with a population grounded in less material than you think it is.
That sounds like a limitation. In practice it’s the most useful design constraint I’ve worked with in a while, because it forces you to decide what each document is for. Every seed document does exactly one of three jobs:
- Population — defines who reacts. The avatar or ICP definition.
- Stimulus — the artifact under test. A sales page, a funnel, a piece of copy.
- Context — grounds the world in reality. Market research, competitive landscape.
One population, one stimulus, one context. If you can’t say which of the three a document is, it doesn’t belong in the run.
The second constraint of the Mirofish system I learned the expensive way: keep the population pure. My two ICPs are a non-technical operations manager and a P&L-owning director. My first instinct was to put both in one project. That was wrong — you get a mixed population debating across two audiences who would never be in the same room, and the conversion signal gets contaminated. The director’s objections to a €697 engagement muddy the operator’s reaction to a €297 course, and you learn nothing clean about either.
Separate projects per ICP for anything offer-specific. Combined populations only when the question is genuinely cross-segment — a funnel that moves one persona toward the other, or a newsletter that serves both.
MiroFish system settings that worked
- Agents: 30–50. Enough for diverse reactions, small enough to stay cheap and stable.
- Rounds: 20–40, starting around 25. Token use grows fast; large runs crash or get expensive.
- Cost: building one pipeline took me roughly 1,200 API requests. In the summer of
2026, that was about 35% of my
deepseek-v4-flashquota, or the wholeglm5.2one. Treat those percentages as a shape rather than a budget — Ollama Cloud doesn’t bill per token or per request; it bills GPU time scaled by how heavy the model is and how long each call runs, so a bigger model charges you twice: once for being bigger and once for being slower. Step 4 dominates either way. It’s roughly one call per agent per round, so cost scales with agents × rounds.
What the AI swarm agents actually found
I ran my operator persona — 40 agents, built from market research and my own avatar definition — and asked whether the segment held together and whether the positioning landed.
The headline finding was one I didn’t want to hear: the avatar overstated technical pain and understated career anxiety.
I had written the persona as someone frustrated that AI tools don’t work reliably. The agents agreed that frustration was real, but it wasn’t what drove them. What drove them was professional exposure — a boss asking about AI with no method to answer, and a peer getting visible credit for being “good with AI.” One agent put it in terms I wouldn’t have written myself:
Agent: “I’m not trying to become an AI expert — I’m trying to become an AI evaluation expert. My literacy is defensive. I’m learning to protect my team and my budget from bad decisions, not to champion the technology.”
The second finding was structural. The population didn’t behave as one segment. It fractured into distinct sub-groups with different trust thresholds — career-climbers who treat AI literacy as a promotion lever, and skeptical practitioners whose posture is, in another agent’s phrase, “not anti-AI, anti-disappointment.”
Those two groups need different copy. The same page cannot speak to both, and I had been writing as though it could.
The part where you should be cautious
The report agent will produce confident, well-written conclusions that are not supported by anything in your documents. This is a known failure mode, and it is the single most important thing to understand before you act on any of this.
Treat the output as a hypothesis generator, not ground truth. The discipline that makes it useful:
- Cross-check every specific claim against your actual seed documents.
- Use the agent-interrogation feature to ask an individual agent “what in the document made you say that?” — a fabricated conclusion falls apart immediately under that question.
- Never let a simulation result override real customer conversations. It ranks what to go ask real people about. It does not replace asking them.
What it’s genuinely good for is surfacing the question you weren’t asking. I would not have thought to test “is career anxiety the real driver?” — I’d have kept optimising copy against the wrong pain. That reframing was worth the setup cost on its own.
What it is not good for is telling you whether your page converts. Only traffic does that.
Would I do it again
Yes, with a clear boundary around what it’s for.
It’s a cheap way to pressure-test positioning before you spend money finding out the expensive way. It’s a genuinely bad way to generate confident claims about your market. The gap between those two uses is where most people are going to get burned with multi-agent tooling over the next couple of years — the output reads like research, and it isn’t.
If you want to try it: start with one ICP, three documents, 30 agents, 25 rounds. Read the report as a list of questions, not a list of answers. Then go and ask a real customer the best question on the list.
None of this is really about agents. I ran this simulation because I’d spent months optimising copy against the wrong pain and had no cheap way to find that out.
That same failure runs inside companies at a much larger scale, and it costs a lot more than a weekend on a virtual machine: months of build aimed at a use case nobody pressure-tested first, and the technology takes the blame afterwards. Choosing the right problem is the part worth being rigorous about, and it’s the part almost nobody schedules time for.
If that’s where you are — a list of plausible AI use cases and no defensible way to rank them — that’s the problem the AI ROI Roadmap is built to solve.