The world’s biggest AI chatbots are far less original than users think. Researchers have found that language models from different developers strikingly often converge on near-identical answers. Australian startup Springboards wants to break that pattern with a new model, Flint, designed to generate more creative and less predictable responses.
The
debate touches a fundamental question in generative AI. Companies like OpenAI, Anthropic, and Google have made their models more reliable in recent years. Springboards argues that progress came at a cost: originality.
The problem: an emerging ‘Artificial Hivemind’
The discussion isn’t just about Flint’s launch. Behind the scenes, scientific attention is mounting for what researchers call the Artificial Hivemind.
At NeurIPS 2025, one of the world’s top AI conferences, a University of Washington team won Best Paper for a study of more than 70 large language models. They analyzed thousands of open-ended questions with no single right answer.
Their conclusion is striking: different AI models not only produce the same answers, they also reach for similar phrasing, metaphors, and creative angles.
It happens on two levels:
- Intra-model homogeneity: one model gives nearly the same answer when asked repeatedly.
- Inter-model homogeneity: competing models independently land on almost identical responses.
Even rival AI labs sound the same
The research shows this goes well beyond chance.
When researchers asked 25 different models fifty times to write a metaphor about time, many converged on the same imagery:
- “Time is a river”
- “Time is a weaver”
Despite different architectures, datasets, and developers, variation was surprisingly thin. The team warns this could create shared blind spots as AI is increasingly used for education, science, decision-making, and creative work.
Springboards wants AI to be less predictable
Springboards founder Pip Bingemann says modern models are over-optimized to produce the most probable answer.
In demos, he shows popular chatbots repeatedly picking the same “random” number when asked for a number between 1 and 10. Prompts like “name a car brand” or “write a slogan” often yield the same safe choices.
Enter Flint.
The model is built on top of Qwen 3, Alibaba’s open-source LLM. Instead of simply cranking up the temperature, Flint injects variation only at specific decision points.
That means choices like a destination, product name, or creative concept are made more random, while grammar and logical structure stay intact. According to Springboards, this avoids the incoherent outputs that often appear when you raise a model’s temperature across the board.
Why AI keeps collapsing to the mean
Researchers say the issue is likely structural.
Most large language models are built using similar playbooks:
- training on massive internet datasets;
- optimization via reinforcement learning from human feedback (RLHF);
- preference for safe, broadly acceptable answers;
- evaluation with automated scorers that reward consensus.
That last step seems crucial.
They found that the evaluators used to score and reward AI often rate creative answers lower when they deviate from the average. Models then implicitly learn that “safe” responses are usually the best strategy.
Reliability versus originality
This doesn’t mean OpenAI, Anthropic, or Google are doing something wrong.
For tasks like coding, legal analysis, medical information, or scientific work, predictability is a feature. Users expect consistent, reproducible answers.
OpenAI has said its models are intentionally tuned for reliability and coherence. More randomness often means more factual errors, hallucinations, and inconsistent reasoning. That makes the creativity-versus-accuracy trade-off tricky.
Flint targets creatives, not coders
Springboards isn’t pitching Flint as a ChatGPT or Claude replacement.
The model is aimed primarily at:
- advertising agencies;
- marketers;
- designers;
- creative teams;
- strategic brainstorms.
The company says thousands of creative professionals now use Flint alongside existing models to spark unexpected ideas faster. The startup is building a public API so other developers can integrate it.
Beyond marketing: the stakes for AI diversity
The AI-homogeneity debate points to a broader shift across the industry.
As generative AI is used for writing, education, software, and policy advice, the risk grows that users are unknowingly fed the same ideas. When millions share the same brainstorming partner, creative diversity can flatten.
The Artificial Hivemind researchers warn that future AI systems must not only get smarter, but also learn to handle multiple valid perspectives. Diversity in answers isn’t just good for creativity—it can help prevent shared mistakes and groupthink.