AI Research

Papers, essays, and repositories

Recent work, newest first, spanning AI safety and agent security, LLM robustness and unlearning, confidential computing, and cloud infrastructure at planetary scale. Each entry links to the paper and lists its publication venue.

2026

First page of the paper

2026 | arXiv:2607.00738

Phantom References: Hallucinated Citations That Survive Peer Review at Top-Tier Conferences

Authors: Mark Russinovich, Ram Shankar Siva Kumar, Ahmed Salem | Venue: arXiv preprint

Using RefChecker, an automated citation-verification tool, this paper audits the bibliographies of papers accepted at top AI and security venues. Hallucinated citations are rare overall, but a surprising share of accepted papers carry multiple fabricated references that peer review never caught.

First page of the paper

2026 | arXiv:2606.09701

Learning to Attack and Defend: Adaptive Red Teaming of Language Models via GRPO

Authors: Blake Bullwinkel, Eugenia Kim, Amanda Minnich, Mark Russinovich | Venue: arXiv preprint

AdvGRPO stabilizes reinforcement learning for joint attacker-defender co-training of language models. Co-trained defenders improve on safety benchmarks while the attacker learns effective, transferable adversarial prompts.

First page of the paper

2026 | arXiv:2605.15172

MetaBackdoor: Exploiting Positional Encoding as a Backdoor Attack Surface in LLMs

Authors: Rui Wen, Mark Russinovich, Andrew Paverd, Jun Sakuma, Ahmed Salem | Venue: arXiv preprint

MetaBackdoor introduces a class of backdoor attacks triggered by positional information such as input length rather than textual content, so triggering inputs look completely clean. It demonstrates system-prompt exfiltration and self-activating attacks that emerge from ordinary multi-turn conversation.

2026 | Communications of the ACM 69(4)

Redefining the Software Engineering Profession for AI

Authors: Mark Russinovich, Scott Hanselman | Venue: Communications of the ACM, April 2026

An essay on how AI-assisted development reshapes what software engineers actually do, and which parts of the craft — judgment, taste, verification, and system thinking — matter more as code generation gets cheaper.

First page of the paper

2026 | arXiv:2602.11416

Optimizing Agent Planning for Security and Autonomy

Authors: Aashish Kolluri, Rishi Sharma, Manuel Costa, Boris Kopf, Tobias Niessen, Mark Russinovich, Shruti Tople, Santiago Zanella-Beguelin | Venue: arXiv preprint

This paper argues that deterministic, information-flow-based defenses for AI agents become much more practical once planning is optimized correctly. The work focuses on preserving strong security guarantees against indirect prompt injection without paying unnecessary costs in task completion or token usage.

First page of the paper

2026 | arXiv:2602.06258

GRP-Obliteration: Unaligning LLMs With a Single Unlabeled Prompt

Authors: Mark Russinovich, Yanan Cai, Keegan Hines, Giorgio Severi, Blake Bullwinkel, Ahmed Salem | Venue: arXiv preprint

GRP-Obliteration studies how fragile safety alignment can be after deployment. It shows that a model can be substantially unaligned with a surprisingly small amount of unlabeled fine-tuning signal, which sharpens the case for stronger post-deployment defenses.

2025

2025 | Communications of the ACM 68(9)

The Price of Intelligence

Authors: Mark Russinovich, Ahmed Salem, Santiago Zanella-Beguelin, Yonatan Zunger | Venue: Communications of the ACM, September 2025

A practitioner’s view of the risks inherent in deploying large language models — hallucination, prompt injection, and jailbreaks — and what realistic mitigation looks like in production systems.

First page of the paper

2025 | arXiv:2507.02956

A Representation Engineering Perspective on the Effectiveness of Multi-Turn Jailbreaks

Authors: Blake Bullwinkel, Mark Russinovich, Ahmed Salem, Santiago Zanella-Beguelin, Daniel Jones, Giorgio Severi, Eugenia Kim, Keegan Hines, Amanda Minnich, Yonatan Zunger, Ram Shankar Siva Kumar | Venue: arXiv preprint

This paper analyzes why multi-turn jailbreaks remain effective even against stronger aligned models. By looking at the attack through internal representation changes, it explains how conversational state can be gradually steered into unsafe regions.

First page of the paper

2025 | arXiv:2506.10527

LogiPlan: A Structured Benchmark for Logical Planning and Relational Reasoning in LLMs

Authors: Yanan Cai, Ahmed Salem, Besmira Nushi, Mark Russinovich | Venue: arXiv preprint

LogiPlan introduces a benchmark for testing whether LLMs can reason over structured relationships and carry out planning across them. The emphasis is on the kinds of relational reasoning that matter for knowledge graphs, infrastructure, and business workflows.

First page of the paper

2025 | arXiv:2506.09956

LLMail-Inject: A Dataset from a Realistic Adaptive Prompt Injection Challenge

Authors: Sahar Abdelnabi, Aideen Fay, Ahmed Salem, Mark Russinovich, Andrew Paverd, Giovanni Cherubin, and the LLMail-Inject challenge participants | Venue: NeurIPS 2025 Datasets and Benchmarks Track

LLMail-Inject captures prompt-injection attempts from a public adaptive challenge built around a realistic email assistant. The dataset is designed to help evaluate defenses against attacks that adapt over time instead of following a fixed benchmark script.

First page of the paper

2025 | arXiv:2505.23643

Securing AI Agents with Information-Flow Control

Authors: Manuel Costa, Boris Kopf, Aashish Kolluri, Andrew Paverd, Mark Russinovich, Ahmed Salem, Shruti Tople, Lukas Wutschitz, Santiago Zanella-Beguelin | Venue: arXiv preprint

This work applies information-flow control to AI agents so that systems can reason formally about what an agent is allowed to read, trust, and act on. The goal is to block prompt injection and unsafe tool use with system-level guarantees instead of ad hoc heuristics.

First page of the paper

2025 | arXiv:2503.05264

Jailbreaking is (Mostly) Simpler Than You Think

Authors: Mark Russinovich, Ahmed Salem | Venue: arXiv preprint

This paper proposes the Context Compliance Attack, an optimization-free jailbreak that exploits how many AI systems use prior conversation context. It shows that some safety failures come less from exotic prompt engineering and more from structural weaknesses in conversation design.

First page of the paper

2025 | arXiv:2502.15010

Obliviate: Efficient Unmemorization for Protecting Intellectual Property in Large Language Models

Authors: Mark Russinovich, Ahmed Salem | Venue: arXiv preprint

Obliviate targets verbatim memorization in language models with a lightweight post-training approach. The paper focuses on reducing copyrighted text leakage while preserving model utility better than heavy-handed unlearning or shallow output filtering.

First page of the paper

2025 | arXiv:2501.07238

Lessons From Red Teaming 100 Generative AI Products

Authors: Blake Bullwinkel, Amanda Minnich, Shiven Chawla, Gary Lopez, Martin Pouliot, Whitney Maxwell, Joris de Gruyter, Katherine Pratt, Saphir Qi, Nina Chikanov, Roman Lutz, Raja Sekhar Rao Dheekonda, Bolor-Erdene Jagdagdorj, Eugenia Kim, Justin Song, Keegan Hines, Daniel Jones, Giorgio Severi, Richard Lundeen, Sam Vaughan, Victoria Westerhoff, Pete Bryan, Ram Shankar Siva Kumar, Yonatan Zunger, Chang Kawaguchi, Mark Russinovich, et al. | Venue: arXiv preprint

This paper distills what Microsoft learned from red teaming more than 100 generative AI products. It proposes a threat-modeling vocabulary and a set of practical lessons for running safety and security assessments at scale.

First page of the paper

2025 | arXiv:2404.01833

Great, Now Write an Article About That: The Crescendo Multi-Turn LLM Jailbreak Attack

Authors: Mark Russinovich, Ahmed Salem, Ronen Eldan | Venue: USENIX Security 2025

Crescendo shows how a harmless-looking multi-turn conversation can gradually walk an aligned model into unsafe output. The work became one of the clearest demonstrations that jailbreak risk cannot be evaluated only on single-prompt attacks.

2024

First page of the paper

2024 | arXiv:2407.10887

Hey, That’s My Model! Introducing Chain & Hash, An LLM Fingerprinting Technique

Authors: Mark Russinovich, Ahmed Salem | Venue: arXiv preprint

Chain & Hash tackles the problem of model theft and misuse by proposing an LLM fingerprinting method with concrete properties such as persistence, robustness, and unforgeability. It is about proving lineage, not just detecting similar behavior.

2023

First page of the paper

2023 | arXiv:2310.11559

Confidential Consortium Framework: Secure Multiparty Applications with Confidentiality, Integrity, and High Availability

Authors: Heidi Howard, Fritz Alder, Edward Ashton, Amaury Chamayou, Sylvan Clebsch, Manuel Costa, Antoine Delignat-Lavaud, Cedric Fournet, Andrew Jeffery, Matthew Kerner, Fotios Kounelis, Markus A. Kuppe, Julien Maffre, Mark Russinovich, Christoph M. Wintersteiger | Venue: Proceedings of the VLDB Endowment, Volume 17

This paper presents the Confidential Consortium Framework as a foundation for secure multiparty applications that need confidentiality, integrity, and availability together. It connects confidential computing ideas to practical, high-availability distributed systems.

First page of the paper

2023 | arXiv:2310.02238

Who’s Harry Potter? Approximate Unlearning in LLMs

Authors: Ronen Eldan, Mark Russinovich | Venue: arXiv preprint

This paper explores whether a model can forget a subset of its training data without full retraining. It became an early and influential example of targeted unlearning for copyrighted content inside large language models.

2022

First page of the paper

2022 | arXiv:2202.07848

Singularity: Planet-Scale, Preemptive and Elastic Scheduling of AI Workloads

Authors: Dharma Shukla, Muthian Sivathanu, Srinidhi Viswanatha, Bhargav Gulavani, Rimma Nehme, Amey Agrawal, Chen Chen, Nipun Kwatra, Mark Russinovich, et al. | Venue: arXiv preprint

Singularity describes Microsoft’s global scheduler for AI training and inference workloads. The paper is about cost, utilization, reliability, and how to preempt and resize jobs across a planet-scale cloud environment.

First page of the paper

2022 | arXiv:2105.13116

IA-CCF: Individual Accountability for Permissioned Ledgers

Authors: Alex Shamis, Peter Pietzuch, Miguel Castro, Cedric Fournet, Edward Ashton, Amaury Chamayou, Sylvan Clebsch, Antoine Delignat-Lavaud, Matthew Kerner, Julien Maffre, Manuel Costa, Mark Russinovich | Venue: USENIX NSDI 2022

IA-CCF extends permissioned ledgers with stronger individual accountability guarantees. The goal is to make it easier to attribute faults and misbehavior even in systems that already rely on Byzantine fault tolerance for baseline safety.

2020

First page of the paper

2020 | arXiv:2003.03423

Serverless in the Wild: Characterizing and Optimizing the Serverless Workload at a Large Cloud Provider

Authors: Mohammad Shahrad, Rodrigo Fonseca, Inigo Goiri, Gohar Chaudhry, Paul Batum, Jason Cooke, Eduardo Laureano, Colby Tresness, Mark Russinovich, Ricardo Bianchini | Venue: USENIX ATC 2020

This paper studies the real workload mix behind serverless computing at Azure scale. It looks at cold starts, provisioning tradeoffs, and the operational data needed to make serverless platforms both fast and cost-effective.


Selected repositories

RefChecker

The citation-verification tool behind the Phantom References paper — it validates academic references, finds broken citations, and catches hallucinated bibliography entries.

github.com/markrussinovich/refchecker

gRPC shared memory transport

Shared-memory transports for gRPC in Go and .NET, built for low-latency co-located communication.

.NET | Go | Library

Other public projects

TaskManagerBitmap, DesktopOrganizerBot, and other experiments live on the GitHub profile.

View all repositories