Came across an article yesterday where someone (IMHO very wrongly) argued that LLMs, because they predict the most likel...

Came across an article yesterday where someone (IMHO very wrongly) argued that LLMs, because they predict the most likely next token, only can create boring blog entries, because suspense implies using unlikely continuations.I believe this point was valid for GPT 2, maybe 3, but in a modern LLM the likely next tokens for "write a story with unpredictable twists" + story beginning will include unpredictable twists. So basically this argument is the dead stochastic parrot in disguise, focusing on how the LLM produces and not how it represents concepts to produce something.Had to try it out with his example, the famous "The quick brown fox jumps over the".So I asked the LLMs to come up with unlikely continuations. All did reasonably well. But the totally crazy one was Mistral.Mistral went into two minutes of deep thinking about my request. Thought things like "Banana? No, too random".And after over 2 minutes it answered. See screenshot.#llm #mistral

Read Original

Related