PERSON DIRECTORY
Nicholas Carlini
Nicholas Carlini appears in 2 indexed conversations across Machine Learning Street Talk, The Cognitive Revolution. This directory brings every appearance, source, TL;DR, digest, and transcript into one searchable feed.
The Adversarial Mind: Defeating AI Defenses with Nicholas Carlini of Google DeepMind
Erik TorenbergNathan LabenzNicholas Carlini
AI defenses often deliver machine-learning accuracy rather than security-grade reliability, with adversarial training retaining roughly 50%–70% accuracy against its training attack class.Attackers move second, inspect the deployed system, and often defeat defenses that merely make gradients noisy, saturated, or difficult to optimize; simpler objectives and stronger optimization repeatedly expose the gap.Open-weight safety remains unresolved, making future model capability a release-risk variable while external action constraints, human review, and layered detection offer a fallible defense-in-depth path worth monitoring.
Language Models are "Modelling The World" [Nicholas Carlini]
Nicholas Carlini argues that ML systems should be designed around models that remain “very vulnerable,” because ordinary attackers can still induce arbitrary failures with off-the-shelf code.The same models can lift expert programming productivity by about 50% while reproducing SQL injection and encryption flaws, leaving security debt and unresolved disclosure norms as deployment risks worth monitoring.

