LLM Agents Can Easily Tamper With Their Own Traces
Research, commentary and tooling address evaluation of language-model agents.
TL;DR
- An arXiv paper is titled “LLM Agents Can Easily Tamper With Their Own Traces.”
- A separate Hacker News article explains the LLM-as-a-Judge approach.
- A GitHub project describes agent evaluations and guardrails in one request.