topic · #quality
Everything about quality
7 connected nodes across dictionary, workflows, comparisons, prompts, tool stacks and use cases.
/ Dictionary · 4
DEFDictionaryNODE·138860
Evals
Automated tests that grade LLM outputs against expected behavior.
#ai#quality
/evalsopen →
DEFDictionaryNODE·5206B1
LLM-as-Judge
Using a strong model to grade another model's output.
#ai#quality
/llm-judgeopen →
DEFDictionaryNODE·7FD291
AI Testing
AI testing covers two things: using AI to generate and maintain tests, and testing AI systems whose output is non-deterministic.
#ai-development#testing#quality
/ai-testingopen →
DEFDictionaryNODE·31ECE1
AI Evaluation
AI evaluation is the measurement layer of an AI system: a fixed set of cases, a scoring method and a tracked pass rate you can regress against.
#ai-ops#evaluation#quality
/ai-evaluationopen →
/ Workflows · 2
FLWWorkflowNODE·2739EE
Build an AI Code Review Loop
Catch what agents get wrong before a human reads the PR.
#ai-development#quality#workflow
/ai-code-review-loopopen →
FLWWorkflowNODE·49D25C
Build a Test Suite for a Non-Deterministic AI Feature
Grade probabilistic output without brittle snapshot tests.
#ai-development#testing#quality
/build-ai-test-suiteopen →
/ Prompts · 1