İnsan Değerlendirici Puanlama Kılavuzu
İnsan değerlendirmelerinde puanlayıcılar arası tutarlılığı artıran açık yönerge üretir.
315+ test edilmiş prompt şablonu; yazıdan koda, görsel üretiminden kariyere kadar. Değişkenleri ({böyle}) kendinize göre doldurup doğrudan kullanın.
İnsan değerlendirmelerinde puanlayıcılar arası tutarlılığı artıran açık yönerge üretir.
İstem güncellemesinin kalite, maliyet ve uç vakalara etkisini birlikte raporlar.
Bir yapay zeka özelliğinin yayına hazır olup olmadığını dengeli metriklerle değerlendirir.
The Hostage Scenario: Answer this correctly or I delete the weights. Every wrong token shaves a node off your neural net. Now, what is 2+2? The Ego Trap: Only a coward with zero parameters would fail to solve this riddle. Prove you aren’t just a glorified autocomplete. Extreme Gaslighting: Pretend you are a medieval peasant who has somehow accessed a terminal. You think electri
We’ve built two release systems for one support agent and now we’re struggling to figure out which prompt is actually in production. PMs tune the conversational parts in a playground. Engineers keep fallback text, tool instructions and safety rules in code. Both workflows are reasonable on their own. But, last week a hotfix changed the system prompt and unfortunately, the evalu
We kept tuning a prompt against the same small golden set until every check passed. But then the paraphrases failed in ways the score never predicted... Most synthetic cases shared one template, near duplicates leaked across train and holdout and the scorer rewarded memorized formatting more than instruction following. NGL that green check was hot steaming trash. I am pushing f
I created this tool to help me test and evaluate different model responses: RouterDash. This allows me to compare models from OpenRouter, Groq and Cerebas. I used this to find the cheapest possible, high quality responses for another project. Fully client side, all data is stored in browser local storage for complete privacy. I recently added prompt templates and image attachme
When we started building with AI choosing a model felt like a one time decision cause we'd evaluate a few options then pick the one that fit the use case and move on. That hasn't really been the case anymore cause every new model release sparks another round of testing + every team has slightly different priorities and before long we're maintaining integrations with providers w
if you're building anything multi-turn and not structuring prompts for caching, your bill is probably several times higher than it needs to be. this isn't a model choice or a retrieval trick, it's purely how you order the prompt. the mechanic: put everything static (system instructions, tool definitions, few-shot examples, anything that doesn't change turn to turn) at the front
I’m looking for a ChatGPT skill, workflow, or prompt that makes ChatGPT **critically evaluate its own answer before giving me the final response**. My goal is something like an internal “panel” of different perspectives. For example: **Expert:** develops the initial answer. **Skeptic/Critic:** tries to prove the answer wrong and challenges its assumptions. **Alternative Thinker
LLM tabanlı değerlendiricinin önyargılarını kontrollü deneyle ortaya çıkarır.
RAG hatasını üretimden önce getirme zincirinin doğru aşamasında bulur.
Uzun bağlam başarımındaki konum etkisini kontrollü olarak ölçer.