arXiv Artificial Intelligence

Persistent Negatives for Adversarial Black-Box On-Policy Distillation

Persistent Negatives for Adversarial Black-Box On-Policy Distillation

Quick summary

arXiv:2609.30864v1 Announce Type: cross Abstract: Black-box On-Policy Distillation (OPD) seeks to improve a student from its own generations when the teacher provides sampled responses but not token probabilities. Adversarial distillation offers one route: it learns a discriminator over prompt-matched teacher and student responses and uses its score as the policy reward. However, sampling discriminator negatives from the latest student at each step couples the learned reward to a negative distribution that changes after every policy update. We address this moving-target problem with persistent

Key takeaways

  • arXiv:2609.30864v1 Announce Type: cross Abstract: Black-box On-Policy Distillation (OPD) seeks to improve a student from its own generations when the teacher provides sampled responses but not token probabilities.
  • Adversarial distillation offers one route: it learns a discriminator over prompt-matched teacher and student responses and uses its score as the policy reward.
  • However, sampling discriminator negatives from the latest student at each step couples the learned reward to a negative distribution that changes after every policy update.

Why it matters

“Persistent Negatives for Adversarial Black-Box On-Policy Distillation” may affect what data AI products can use and where accountability sits. Product teams should watch compliance duties, rights holders should watch enforcement, and users should watch transparency and appeal mechanisms.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗