Why 'Don't Be Generic' Doesn't Work
By Lovro Lucic ·
The "It Depends" Problem · 3 of 5
Asked for a competitive analysis. Gave it everything. Market position, three-year data, the specific situation. "Be insightful. Don't be generic."
The output was structured, fluent, professional. And interchangeable with what it would have produced for any company in any market. Every recommendation could have been copy-pasted into a competitor's strategy doc without changing a word.
Twenty controlled runs confirmed what was visible, a specificity-by-framing study separate from the receipt-backed 2x2 below. Specific versus vague crossed with positive versus negative framing. Four combinations. Pure negation ("don't be generic," "avoid clichés," "don't use buzzwords") was indistinguishable from giving no instruction at all.
That removes a label without providing a destination. The model left the default and wandered to the adjacent region. Same neighborhood. Different house number. This is the specificity case of that larger pattern: negation names what to avoid, which the default already routes around; only an anchor forces a different path.
Then the other cell in the design. Same negation, but paired: "Don't include recommendations that could apply to any B2B SaaS company. Every recommendation must reference Northvane's specific assets, 5 years of shipping logistics data, 12 engineers, Pacific Northwest enterprise incumbents."
Strong effect, replicated in direction across generators. An earlier comparison looked Claude-leading, but it was confounded by instruction length and is not comparable; a clean 2x2 on xAI, matched on Gemini Flash, shows the effect is not Claude-specific (raw-versus-density magnitudes in the footer). This is the content-specificity lever. Stacking more explicit constraints is a different lever from a different experiment, model-specific rather than cross-generator: large on GPT, reversed on Gemini, uninformative on Claude. Keep the two separate.
The specificity isn't measurably adding analytical quality. It's adding verifiability.
"Don't be generic" blocks one path. The model takes the next most likely path, which is a variation on generic. "Reference these specific assets" creates an anchor the output has to pass through. The result physically cannot be the same for a different company. The constraint tests itself.
Here's the nuance that matters: in blind testing (5 pairs, domain expert, evaluator's own expertise area), the expert couldn't distinguish specific from generic outputs on quality. Picked specific 3 out of 5 times: chance level. Five pairs is too few to prove the conclusions identical; what it shows is no detectable quality difference. The specificity instruction changes what the output LOOKS LIKE (more data references, more grounded claims), not what it SAYS, at least not measurably.
The demonstrated value is verifiability. The specific output can be checked: every claim traces to something nameable. The generic output makes the same points but you can't verify them. When you need to trust the analysis (regulatory, audit, high-stakes decisions), specificity makes the output auditable. When you're using it as a starting point for your own thinking, it doesn't matter.
Test this yourself
Next time you write "don't be X," finish the sentence: "instead, Y with Z criteria." Run both versions. Measure the difference.
What survived testing
- Negation alone has no effect. Specificity is real, and the specificity effect replicates across generators. (That specific-plus-negative scored highest is the framing experiment's ranking, not itself cross-validated.) The clean magnitude is g=1.34 on raw marker count and 1.62 at density on xAI, nearly identical at density on Gemini Flash (g=1.64) though not on raw (g=0.65); the Claude-versus-others differences came from a length-confounded comparison and are not clean. Quality demands are density-dependent. Synergistic with specificity, not additive (quality demands do little alone, but add on top of specificity, which already works; together the effect is much larger). Negation removes label without destination. Specificity provides anchor.Copy link
What didn't survive
- "Negation hurts" overclaimed. Negation is null, not negative. "More specific = better" linearly is unverified: the effect was measured as specificity present versus absent, not across a density gradient, so behavior at extreme specificity is untested. "At density is the pure confound-free measure" too strong: density (markers per 1,000 words) removes the prompt-length confound but partially conflates specificity with output brevity, because shorter outputs score higher density. It strips one confound, not all. So the cross-generator match is on density; the replication is directional, not a magnitude match.Copy link
Honest limits
- Single operator. Transfer to other operators untested.Copy link
- The negation result and the specificity-magnitude result come from two different experiments: a specificity-by-framing design shows negation alone is null, while the receipt-backed specificity-by-quality-demands 2x2 establishes the magnitude and the cross-generator replication. Only the second is in the bound receipt; the negation-by-framing result is not reader-auditable there.Copy link
- The constraint-count lever's model differences are not all clean reads: Claude's 0.00 is a rubric ceiling (25/25, zero variance in both conditions), not an absence; the GPT result is self-scored (the model graded its own outputs), so treat it as an upper bound.Copy link
- Effect sizes measure programmatic specificity markers (company mentions, scenario numbers, market terms). These counts are objective. Domain expert validation (5 blind pairs, evaluator's own domain): indistinguishable from chance. Expert rated based on style (rhythm, naturalness), not on marker density. At that sample size no quality difference was detectable, which is not proof there is none. Specificity demonstrably changes output FORM (more verifiable references); a SUBSTANCE difference was undetectable here, neither shown nor ruled out. What is demonstrated is verifiability.Copy link
Next in The "It Depends" Problem
More Context Barely HelpsExplore other threads
The Fabrication Problem
5 findingsMost AI numbers are fabricated. Source material fixes it. Self-checking fails. Trust signals are backwards.
The Evaluation Problem
2 findingsJudgment goes quiet. You can't see the gaps. Satisfaction is the trap. Stronger evaluators discriminate less.
The "What You Think Works" Problem
1 findingTemporal decay is a myth. Self-critique circles. Constraints narrow. Quality ceiling per mode.
New findings when they land.
No spam. Just what held up.