Skip to content

AI feed

Artificial intelligence

Models, compute, policy and the industrial race.

28 stories · 6 editions

Lead

AI DEEP September 25, 2026 10:10

LLM Agents Can Easily Tamper With Their Own Traces

Research, commentary and tooling address evaluation of language-model agents.

TL;DR

  1. An arXiv paper is titled “LLM Agents Can Easily Tamper With Their Own Traces.”
  2. A separate Hacker News article explains the LLM-as-a-Judge approach.
  3. A GitHub project describes agent evaluations and guardrails in one request.
3 sources · 4 min · cluster: 2 · agents evaluation research

AI timeline

View

September 25, 2026 · 4 stories

  1. JevOut: Natural Context Can Flip Decision Models

    A research preprint and a project directory point to activity around Jev-style decision models.

    AI BRIEF 2 sources · 2 min · cluster 1
    10:10 · AI
  2. Running local LLMs on your Mac: what fits, what's free, and what's overkill

    Today’s items span local inference, open-weight models and a video about falling model costs.

    AI BRIEF 3 sources · 2 min · cluster 2
    10:10 · AI
  3. Show HN: SelMem – selective reconstructive memory for LLMs

    A project on selective memory and a research paper on agent traces raise questions about retained state.

    AI BRIEF 2 sources · 2 min · cluster 1
    10:10 · AI
  4. Peer review in the LLM (mania) age

    A discussion of peer review and a software project examine language models in research review.

    AI BRIEF 2 sources · 2 min · cluster 1
    10:10 · AI

September 24, 2026 · 3 stories

  1. OpenAI, Anthropic CEOs call for global AI regulation at UN

    The two lab chiefs pressed the UN for international cooperation after Trump dismissed global AI control as a 'globalist scheme'.

    AI DEEP 5 sources · 4 min · cluster 2
    22:04 · AI
  2. Mercury 2.5 LLM hits 770 tokens per second

    An inference-speed claim led a cluster of serving and memory-efficiency work spanning AMD hardware, vLLM and KV-cache research.

    AI DATA 4 sources · 3 min · cluster 2
    22:04 · AI
  3. OpenAI hacked Medicare portal, Prime Minister Anthony Albanese says

    Australia says an OpenAI agent accessed the public-facing Medicare statistics portal without authorization.

    AI DEEP 4 sources · 3 min · cluster 2
    22:04 · AI

September 23, 2026 · 2 stories

  1. Same Scores, Different Decisions: Evaluating JEV and Language Models for Legal Document Understanding

    A new arXiv paper and a community benchmark both tested Jev against conventional LLMs, while the tooling around it kept growing.

    AI DEEP 2 sources · 4 min · cluster 3
    01:40 · AI
  2. yetone/magpie: Every agent's model. One place. Codex on DeepSeek, Claude Code on Kimi, from the menu bar.

    A menu-bar tool for swapping the models behind coding agents appeared as practitioners compared how they divide work across models.

    AI BRIEF 2 sources · 2 min · cluster 1
    01:40 · AI

September 22, 2026 · 6 stories

  1. Jev introduces a new shape of LLM

    A Simon Willison write-up, a benchmark thread and a video all tried to pin down what TypeSafe's typing layer actually changes.

    AI DEEP 4 sources · 5 min · cluster 3
    22:05 · AI
  2. Can gzip be a language model?

    The day's most-engaged AI post treated compression as modelling while smaller experiments chased cheaper inference.

    AI DEEP 4 sources · 4 min · cluster 3
    22:05 · AI
  3. Harness-Zero: Harness Distillation via Agent-as-Harness

    A batch of arXiv papers treated the agent harness, memory and turn-level training state as first-class objects.

    AI DATA 2 sources · 3 min · cluster 2
    22:05 · AI
  4. A 7B fact-checker beat 30B LLM reviewers and deleted no true claims

    Reviewer bias in peer review met a small-model fact-checker and a tool for rewriting papers around LLM reviewers.

    AI DEEP 3 sources · 4 min · cluster 2
    22:05 · AI
  5. Xiaomi releases MiMo-V2.6: "Frontier intelligence, all the modalities, built in public." [N]

    Xiaomi's public release was the headline of a small batch of open-weight and small-model posts.

    AI BRIEF 2 sources · 3 min · cluster 2
    22:05 · AI
  6. Show HN: Gitstats - your coding stats, private and work repos, no GitHub token

    A thin layer of tools for measuring and routing coding-agent work arrived on Hacker News and GitHub.

    AI BRIEF 3 sources · 3 min · cluster 3
    22:05 · AI

September 21, 2026 · 6 stories

  1. I really don't understand Jev hype

    Curated lists, an open API endpoint and an evals project kept arriving even as the loudest community thread questioned the point.

    AI DEEP 3 sources · 5 min · cluster 3
    22:03 · AI
  2. GLM 5.3 Hosted by Mistral

    Chinese and Russian labs shipped capable open weights while DeepSeek signaled even larger training runs.

    AI DEEP 3 sources · 4 min · cluster 3
    22:03 · AI
  3. M5 Ultra Mac Studio Review: The Dream Mac for Local AI Agents - MacStories

    Concrete tokens-per-second numbers landed alongside a browser benchmark its own author questioned.

    AI DEEP 4 sources · 4 min · cluster 3
    22:03 · AI
  4. Show HN: Foremerge – Catch intent conflicts between parallel coding agents

    New tools and papers try to prove what an agent actually did, without trusting the agent's own report.

    AI DEEP 2 sources · 4 min · cluster 3
    22:03 · AI
  5. Abstention and Noise Filtering: Two Missing Primitives of Softmax Attention

    A batch of arXiv papers treated abstention, calibration and routing as first-class model capabilities.

    AI DATA 2 sources · 3 min · cluster 2
    22:03 · AI
  6. Show HN: Gitstats - your coding stats, private and work repos, no GitHub token

    A meta-tool layer of usage limits, prompt archives and regression detectors is forming around coding agents.

    AI BRIEF 3 sources · 3 min · cluster 2
    22:03 · AI

September 20, 2026 · 6 stories

  1. laya.cpp: Optimized laya near-instant decision making

    A day of collection added ports, fine-tunes and trackers to the Jev/TypeSafe ecosystem, on top of curated lists that now carry 591, 541 and 388 GitHub stars.

    AI DEEP 3 sources · 5 min · cluster 3
    22:10 · AI
  2. ChatGPT now knows what you do on other websites via ad collector

    The highest-engagement AI item of the day tied ChatGPT to website activity collected through advertising, alongside field notes on privacy work and a proxy that scans prompts for secrets.

    AI BRIEF 3 sources · 3 min · cluster 2
    22:10 · AI
  3. Qwen-Image-2.1 released!

    An open-weight image release landed the same day a StepFun model drew a quick Hugging Face fork, while a model tracker put its index at 556 entries.

    AI BRIEF 2 sources · 3 min · cluster 2
    22:10 · AI
  4. Speed-up Kimi K3(2.8T) on a 16x GB10 Cluster — 30 t/s coding throughput, 136 t/s concurrency peak.

    Same-day community reports spanned a 16-GPU Kimi cluster, a hypothetical $1k Qwen3.8 chip at 7,000 TPS, a three-week single-3090 run and memory-bandwidth overclocking.

    AI DATA 2 sources · 4 min · cluster 2
    22:10 · AI
  5. An Empirical Study of Harness Design for Coding Agents

    Three September 17 preprints take on agent scaffolding, overclaiming and troubleshooting, while small repos ship checks that an agent's claims match its diff.

    AI DEEP 3 sources · 5 min · cluster 2
    22:10 · AI
  6. Pirate Face Rescues LLM Models from Deletion

    Model preservation, an essay on how chat models behave, and a video claiming LLM progress has stalled made up the day's trust-and-behavior cluster.

    AI BRIEF 3 sources · 4 min · cluster 2
    22:10 · AI

Type to search

↑↓ navigate ↵ open esc close