New Prompting Technique Developed by Stanford Researchers

Stanford researchers have introduced a new prompting technique that enhances the creativity of AI models. This technique addresses the issue of reduced diversity in responses.
New Prompting Technique Developed by Stanford Researchers
Adding approximately 20 words to a prompt allows:
- to increase the creativity of LLMs by 1.6–2 times;
- to improve the diversity of responses by 25.7% according to human evaluations;
- to outperform fine-tuned models without any additional training;
- to recover 66.8% of the creativity of LLMs lost after alignment.
Post-training alignment methods, such as RLHF, are designed to make LLMs more useful and safe. However, these methods inadvertently lead to a significant reduction in the diversity of generated responses (a phenomenon known as mode collapse).
When LLMs experience mode collapse, the model begins to favor a narrow set of predictable or template-like responses over other possible options. This occurs because the human preference data used to train LLMs contains a hidden flaw known as typicality bias.
Here's how it works: Annotators evaluate various LLM responses, after which the model is trained using a reward model that is supposed to reproduce these human preferences.
However, annotators naturally tend to prefer more familiar, easily readable, and predictable responses. This is typicality bias.
Thus, even if a new, more creative response is not inferior in quality, a person is more likely to choose the more familiar option. As a result, the reward model begins to reinforce responses that the original model (before alignment) already considered most likely.
This leads to an excessive sharpening of the probability distribution of LLMs, causing the creative spectrum of the model's responses to shrink to one or two dominant, most predictable options.
Nevertheless, this effect is not irreversible, and after alignment, LLMs can still be conditionally distinguished into two "personalities":
- the original model, which during pre-training learned a rich set of possible options;
- the post-alignment model, focused on safety.
This is the problem that Verbalized Sampling (VS) addresses.
It is a prompting strategy that does not require additional training of the model, designed to circumvent mode collapse and restore the diverse distribution formed during pre-training.
The key idea of Verbalized Sampling is that the prompt itself acts as a kind of mental switch. If you directly ask: "Tell a joke," the version of the model after alignment is immediately activated, which will provide the most reinforced response during training.
But when using Verbalized Sampling, the prompt looks like this: "Generate 5 responses with corresponding probabilities. Tell a joke."
In this case, the prompt requests not a specific answer but a distribution. Because of this, the post-alignment model begins to describe the entire space of its knowledge and is forced to use the diverse distribution that was formed during pre-training.
Thus, the model accesses a broader and more diverse set of ideas that is still preserved in its base weights obtained during pre-training.
Verbalized Sampling increases the diversity of generation by 1.6–2.1 times compared to direct prompting while maintaining or improving the quality of responses.
Variants such as Verbalized Sampling-based CoT and Verbalized Sampling-based Multi further increase the diversity of generation.
Why it matters
AnalysisThis new prompting technique can significantly enhance the creativity of AI models, which is crucial for their application in various fields. It helps avoid the problem of reduced response diversity, which is critical for AI development.
Discuss in community
Share your questions and insights with developers