Brainstormer

◇ Guide Jul 24, 2026 9 min read

Is AI good at brainstorming? What the research actually shows

By the Brainstormer team

◇ Try it while you read

Pick a challenge, flip the lens, then press cluster and decide. This is the live studio.

Get started
Brainstorm Studio
Try it: pick a challenge

Challenge:

+ more on the wall

Winner

score / 10

Why

Sample brainstorm shown. Your challenges stay private.

Yes, at the parts of brainstorming that are about volume and range. In a Wharton and Cornell study, GPT-4 produced 200 product ideas in about 15 minutes, against roughly five ideas per 15 minutes for one person, and an evaluation panel rated the machine's ideas higher on purchase intent: 47 percent against 40 percent for the MBA students. Where AI is weak is novelty and context, and both are fixable if you know they are there.

That is the honest short answer, and it is worth being clear about who is giving it. We sell an AI brainstorming tool, so treat everything below as an argument you should check. The two studies cited are public working papers, the numbers are theirs, and the section on where AI fails is the part we would rather you read.

Is AI good for brainstorming ideas?

The best available evidence comes from a working paper titled "Using Large Language Models for Idea Generation in Innovation" by Lennart Meincke, Karan Girotra, Gideon Nave, Christian Terwiesch and Karl Ulrich, published through Wharton's Mack Institute in 2023. The design is unusually clean. The researchers took 200 ideas submitted by an MBA class that had been asked for a product aimed at the college student market retailing at $50 or less, then had GPT-4 generate two more pools of 200 against the identical brief, one with a plain prompt and one seeded with example ideas.

Every idea was then put in front of an evaluation panel that estimated purchase probability. The results:

GPT-4 against MBA students on the same product brief (Wharton Mack Institute, 2023)
Idea sourceAverage purchase probabilitySpeed
MBA students40 percentAbout 5 ideas per person per 15 minutes
GPT-4, plain prompt47 percent200 ideas in about 15 minutes
GPT-4, seeded with examples49 percent200 ideas in about 15 minutes

The speed gap is the finding people quote, and it is the less interesting one. The quality gap is what should change how you run a session. A graduate business class working a brief they had time to think about did not beat a model that had the brief for a quarter of an hour, on the specific measure of whether a stranger would buy the thing.

Where AI brainstorming is genuinely weak

The same paper found that the AI ideas were rated less novel than the human ones and showed higher pairwise similarity, meaning they clustered. This is the failure mode that matters, and it is easy to miss because the output looks abundant. Two hundred ideas that are all variations on four underlying concepts is not two hundred ideas. It is four, presented in a way that makes the room feel finished.

A follow-up paper by Meincke, Ethan Mollick and Terwiesch, "Prompting Diverse Ideas: Increasing AI Idea Variance", went after exactly this. They compared prompting methods on how much variance the output actually carried, measured by cosine similarity and by the number of unique ideas before the space got exhausted. Chain-of-thought prompting produced the highest diversity of everything they tested and the largest count of unique ideas, coming close to what groups of human subjects achieve. The lesson is not "AI is or is not diverse". It is that diversity is a property of the prompt, not of the model, and a single naive request buys you the narrow version.

The second weakness is context, and no prompting technique fixes it. A model does not know that your biggest account threatened to leave over exactly this feature in March, that legal has already refused the obvious version twice, or that the person who would have to run the winning idea is leaving in six weeks. Those constraints are what separates an idea that is good in general from one that is good here.

Which AI is best for brainstorming ideas?

For raw generation, the frontier general models are close enough that the choice barely matters: current Claude, GPT and Gemini releases will all produce usable range if you prompt them for it. The differences that actually show up in a working session are about structure rather than the underlying model.

A general chat assistant gives you a wall of text in one thread, fans out from your framing (so your first prompt quietly caps the output), and forgets the session unless you keep pasting it forward. It is excellent for riffing and poor as the backbone of a group session. A purpose-built idea generator runs named techniques over the challenge, tags each idea with the angle it came from, and keeps the wall as an object you can cluster and score. If you want the head-to-head, we wrote it up under using ChatGPT for brainstorming, and the full field including whiteboards is on best brainstorming software.

Can AI have creativity?

Not in the sense of wanting something, and yes in the sense that matters for a workshop. Creativity in idea generation is largely recombination: taking a mechanism from one domain and applying it in another, dropping a constraint everyone treats as fixed, exaggerating one variable until the shape of the problem changes. Models are good at that because they have read across more domains than any person in the room has.

What they do not have is taste or stake. A model will not notice that idea 14 is the one that quietly solves the problem you were too embarrassed to write in the brief. It will not defend an unpopular idea in a meeting, and it does not care whether the project succeeds. That is the division of labor: the machine supplies range, you supply judgment. Every serious use of AI in ideation runs that way, and it is the same argument behind divergent and convergent thinking, which says the widening and the deciding must be separate steps run by separate rules.

Is using AI to brainstorm cheating?

In a work context, no, and the question is already dissolving. Nobody thinks a spreadsheet is cheating at arithmetic. The professional norm has shifted fast in the same direction wherever the model is used to prepare rather than to submit: people now draft with one, pressure-test a strategy with one, and rehearse a salary negotiation with one before the real conversation. The output still has to be yours in the sense that you chose it and you can defend it.

Two places to be careful. Academic settings have their own rules and you should read them rather than assume. And confidential material should not be pasted into a consumer chat account: check whether your tool trains on your inputs and whether your employer has an approved list. That is a procurement question, not an ethics one, but it is the one that actually gets people in trouble.

Is AI brainstorming better than a group brainstorm?

It is better at the part groups are worst at. Decades of research on group ideation keep finding the same losses: production blocking, where only one person can talk at a time, and evaluation apprehension, where people privately discard their odder ideas before saying them. We covered the mechanics in why group brainstorming fails. A model has neither problem, which is why its output is wider than a room's in the first twenty minutes.

Groups are better at everything after that. They know the constraints, they carry the political information, and a decision the room helped make is a decision the room will actually execute. The setup that beats both is sequential: generate wide with the machine before anyone speaks, then let the humans argue, cluster and score. That order also removes the anchoring problem, because nobody's first suggestion gets to define the space.

How to brainstorm with AI without getting the narrow version

Four habits do most of the work.

  • Ask for the reasoning, not just the list. The diversity paper found chain-of-thought prompting produced the widest output of anything tested. Making the model work through the space before it names ideas is the cheapest quality upgrade available.
  • Constrain the angle, not the answer. "Give me 10 ideas" fans out from your framing. "Give me 10 ideas that remove a step rather than add a feature", then "10 that only work if we are wrong about who the customer is", produces genuinely different territory. This is what SCAMPER and Six Thinking Hats automate.
  • Seed it with your own half-formed ideas. The seeded pool scored highest in the original study, at 49 percent. Your bad ideas contain context the model cannot infer.
  • Keep going past the point it feels finished. The first dozen outputs are the obvious ones. The interesting material sits in the second half of a long run, which is the same reason rapid ideation sets a quota rather than a time limit.

Then converge, because none of this counts until something is chosen. Group the wall into named themes, score the themes on impact against effort, and write down why the winner won. That sentence is the only part of the session that survives into next month, and it is the part no model can write for you, because it depends on what you know and what you are willing to defend.

◇ Run it, don't read it

Range from the machine, judgment from you. Put a real challenge in and see what the wall looks like.

Get started