Attention Smoothing Unlearning
A framework that constructs a forget-teacher by smoothing a model’s attention, then uses self-distillation to suppress unwanted factual recall while preserving coherent responses and model utility.
Projects
A project-oriented view of the methods and questions represented in my publication record.
A framework that constructs a forget-teacher by smoothing a model’s attention, then uses self-distillation to suppress unwanted factual recall while preserving coherent responses and model utility.
Research on automatically calibrating membership inference attacks for large language models.
Work investigating token-aware approaches to forgetting in large language models.
Research examining how adversarial demonstrations in context can hijack large language models.
Ongoing work on deceptive safety alignment and counter-aligned few-shot conversation exposure in large reasoning models.
Research on poisoning large language models during training for downstream behavioral manipulation.