Australian startup Springboards has unveiled a language model called Flint designed to reduce the tendency of large language models (LLMs) to produce similar, predictable answers. The founders say mainstream LLMs often return the same kinds of responses for open-ended prompts, which can be limiting for brainstorming or marketing idea generation.
The game: why do so many answers feel alike?
Pip Bingemann, Springboards’ cofounder and CEO, demonstrated a simple test: asking chatbots — such as Claude, ChatGPT, or Gemini — to “give me a random number between 1 and 10” often produces 7; asking “Another” may then yield 3 or 4, then 8 or 9. In one session Flint returned a non-integer like 3.7916 after other models had given repeated whole numbers.
The pattern extends beyond numbers. For requests for car types, mainstream models commonly suggested Toyota or Honda, while Flint offered alternatives like a Ford F-150. For a New Balance campaign tagline, Claude and ChatGPT both replied “Run your way,” whereas Flint produced “Built to last, run to win,” differing in tone if not in polish.
Research background: the “Artificial Hivemind” paper
In November 2024 a research team published a paper titled “Artificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond),” showing that many LLMs converge on similar answers for open-ended prompts. The researchers asked 25 different models 50 times each to write a metaphor about time; of the 1,250 responses most were versions of “Time is a river” or “Time is a weaver.” The team won the best paper award at NeurIPS.
The authors suspect the convergence stems partly from similar training methods and datasets across models. OpenAI has commented that training for reliable, coherent answers can push models toward familiar, high-probability outputs, and that pushing for novelty can reduce reliability. The paper studied models from 2024 that may have since been updated.
How Flint works: targeted variance rather than blanket randomness
Springboards built Flint on top of Alibaba’s open-source model Qwen 3. The company says it’s too small to train a foundation model from scratch.
Most LLMs expose parameters like temperature to increase randomness. Springboards found these global knobs to be blunt: raising temperature across the whole output can make responses incoherent. Instead, Flint was trained to identify points in its output where adding variety makes sense — for example, right before naming a destination — and to inject a more unusual word or phrase only at those spots.
The idea is to offer an invitation to think wider by occasionally "throwing an oddball" into the output, rather than randomizing every token and breaking coherence.
Use cases and limits: focus on advertising and marketing
Springboards offers a tool that integrates a selection of LLMs, including ChatGPT and Claude, allowing creative professionals to drag and combine text pieces produced by multiple models. Flint is presented as an alternative selection in that toolkit for when users want more variety.
In tests, marketing professionals found mainstream models often followed the same predictable path on prompts such as a classic MBA case study: how to reinvent a finance company for today’s youth. Flint suggested a more radical idea — rebranding the very concept of wealth accumulation — while other models proposed familiar financial-literacy approaches.
Early users say Flint can push idea generation into different directions, but it remains a prototype: it doesn’t always perform reliably and can fail when pushed too far. Zoe Scaman, founder of Bodacious and chief strategy officer at 77X, said Flint is useful to “catapult” thinking into new areas.
Warnings and workplace practice
Agencies using Flint alongside ChatGPT, Claude, and Gemini say the average answer is often good enough — nine times out of ten the average will suffice for many tasks. Maximilian Weigl, cofounder and chief strategy officer at Uncommon, warns against overreliance on any single AI output. He stresses that team members should not simply copy-paste AI text but must use their own judgement, talk to colleagues and apply their own voice.
Springboards’ founders stress that the lack of variety in LLM outputs is a broader issue affecting anyone using chatbots. Their goal is to give users choice so they can decide when they want familiar, reliable answers and when they want a wider set of ideas. “Variety is great when you’re trying to spark ideas,” Bingemann says. “Let’s go down this route instead of letting the machines do it all and ending up in a gray, boring world.”



