FlavourBench: Executable Culinary Reward Maps for Language Model Evaluation and Post-Training
Quick summary
arXiv:2608.20574v3 Announce Type: replace Abstract: We introduce FlavorBench: a benchmark for Compiling Dense Deterministic Answer Maps from a Versioned Culinary Embeddings Model. We test 27 frontier large language model endpoints on 534 substitution, pairing and constraining tasks for tasks that request a 3-ingredient portfolio from 8 candidates and score all 56 resulting portfolios. We conducted multiplicity-controlled paired tests on 101 of 351 model contrasts for this task-set. The largest point estimate on this task-set was achieved by Grok 4.6 at 65.1. The same rankings for this task-set
Key takeaways
- arXiv:2608.20574v3 Announce Type: replace Abstract: We introduce FlavorBench: a benchmark for Compiling Dense Deterministic Answer Maps from a Versioned Culinary Embeddings Model.
- We test 27 frontier large language model endpoints on 534 substitution, pairing and constraining tasks for tasks that request a 3-ingredient portfolio from 8 candidates and score all 56 resulting portfolios.
- We conducted multiplicity-controlled paired tests on 101 of 351 model contrasts for this task-set.
Why it matters
“FlavourBench: Executable Culinary Reward Maps for Language Model Evaluation and Post-Training” should be evaluated beyond branding and benchmark scores. Its practical importance will emerge in task accuracy, latency, unit cost, safety and integration with real workflows.

Member comments