🎯
Veri & Analiz

Model Hakemi Kalibrasyon Testi

LLM tabanlı değerlendiricinin önyargılarını kontrollü deneyle ortaya çıkarır.

PROMPT
Bir model hakeminin puanlama güvenilirliğini ölçmek için deney tasarla. Aynı yanıt çiftlerini ters sıra, eşit uzunluk ve gizlenmiş model adı koşullarında yeniden değerlendir. Konum ve uzunluk önyargısını ayrı hesapla; insan uzmanlarla uyuşma oranı için güven aralığı ver. Sonuçta hakemin hangi görevlerde kullanılmaması gerektiğini yaz.
1CUSTOMIZE

Replace the bracketed fields with your own goal, audience and context.

2RUN

Paste the prompt into the recommended tool; treat the first output as a draft.

3REFINE

Point out gaps, add examples, and define the output format you want more precisely.

Recommended tool: ChatGPT

We recommend using this prompt with ChatGPT.

ChatGPT Page →