How to tell whether an LLM system works, and how to keep it working: test cases, graders, judges, groundedness, CI gates, live signals, traces, latency, schemas, injection defense, and PII masking.
An agent that gets better on its own has to change something about itself, and something has to check the change. This post walks the six things it can change, the checker that decides whether any of it worked, the ceiling it runs into, the ways it breaks, and the loop teams actually run in production.
What graph engineering is, how it differs from loops and workflow engines, when to build one, a refund approval worked through end to end, and the failure modes to watch for.