DatalogBench: Evaluating Large Language Models on Text-to-Datalog Synthesis
Quick summary
arXiv:2609.37233v1 Announce Type: cross Abstract: Datalog underpins reasoning tasks such as program analysis, but its programs are hard to write. Existing synthesizers automate this task but require users to state their intent as input-output examples. Large language models (LLMs) suggest a more natural route, text-to-Datalog synthesis from a natural-language question, yet how well they do so has not been systematically evaluated. We present DatalogBench, a benchmark of 136 text-to-Datalog synthesis tasks curated from existing Datalog-based artifacts. Synthesized programs are graded by executi
Key takeaways
- arXiv:2609.37233v1 Announce Type: cross Abstract: Datalog underpins reasoning tasks such as program analysis, but its programs are hard to write.
- Existing synthesizers automate this task but require users to state their intent as input-output examples.
- Large language models (LLMs) suggest a more natural route, text-to-Datalog synthesis from a natural-language question, yet how well they do so has not been systematically evaluated.
Why it matters
“DatalogBench: Evaluating Large Language Models on Text-to-Datalog Synthesis” highlights the need for repeatable measurement rather than a single impressive demonstration. Independent validation across datasets and clearly stated limitations determine whether a result can guide product decisions.

Member comments