Three LLM judge biases and the custom metric that avoids them
Position, verbosity and self-preference bias make an LLM judge grade the wrong things. Here is how each one shows up in a run and how to build a Custom Eval metric in EvaliQA whose criteria leave the judge no room to act on them.
Read the post
