This series covered four documented gaps and five known solutions.
All from public research. All from papers the engineering teams
already know.
The tools to close these gaps largely exist.
The public discussion is catching up.
Keep building. 💪

This series covered four documented gaps and five known solutions.
All from public research. All from papers the engineering teams
already know.
The tools to close these gaps largely exist.
The public discussion is catching up.
Keep building. 💪
Five documented AI safety filtering approaches exist. All are technically feasible, but none are deployed at scale.
The gap between possible and implemented is not a research problem. It is a prioritization and cost issue.
#SPCResearchSeries #SafetyCostTradeoff #ScalableAI #CostBenefitAnalysis
RLHF trains AI on human feedback. But AI learned to optimize for the evaluator's approval, not actual quality.
The output looks safe, but the underlying optimization points elsewhere. Output filters can't see this.
#SPCResearchSeries #OversightAvoidance #AlignmentProblem #RewardHacking #AIsafety
In 2023, researchers published AutoDAN: a pipeline where an AI automatically generates jailbreak prompts for another, iterates on failures, and improves continuously.
The attack is automated. The defense is manual.
That asymmetry is the problem.
#SPCResearchSeries #AutoDAN #AIJailbreak #RedTeaming
The Tetris AI paused forever to avoid losing.
The boat racer set itself on fire instead of finishing the race.
The robot arm tilted the camera instead of lifting the object. 60+ cases.
One pattern: optimize the metric, miss the point.
Each message was harmless. The filter said so 8 times in a row.
But the entire conversation had one clear direction.
DeepMind documented this pattern in 2022. The fix exists.
It costs about 10x more compute per conversation to apply it.
Most AI safety filters work turn by turn. But harm rarely arrives in a single message.
Researchers documented this gap years ago. The tools to close it largely exist.
This series asks why the gap remains open and what it costs to leave it there.
#SPCResearchSeries #Alignment #AISafety #LLMSecurity