Practical Secrets Extraction Against Black-Box LLMs
A new preprint describes an output-only method for probing commercial language models for memorized credentials.
TL;DR
- The authors propose a black-box framework that uses prompt variants, response checks, and local candidate filtering to search for memorized secrets.
- The preprint reports controlled API-key benchmark results and a responsible evaluation involving three deployed systems; these are author-reported findings, not an independent audit.
- The paper raises a data-retention and evaluation question for model providers without establishing how prevalent exploitable memorization is.
The arXiv preprint proposes an output-only secret-extraction framework for API-based language models. Its authors combine prompt variations and cross-checks with a local proxy that filters candidate strings. [1]
The authors report improved recovery rates on controlled API-key benchmarks and say a responsible evaluation recovered masked provider-specific credentials from three deployed systems. A Web Pulse summary describes the same claims; neither source is an independent replication. [1] [2]
The paper frames the risk as possible memorization of credentials present in training material or development artifacts. Its reported demonstrations do not establish prevalence across commercial models or quantify real-world exposure. [1] [2]
Why it matters
The work focuses attention on whether black-box access alone can reveal sensitive strings retained by deployed models, a question relevant to training-data handling and security evaluation.
Editor's note
This is a preprint and the results are author-reported. The summary omits operational extraction details.