1,485 concepts
AI agents coordinating by leaving traces in a shared environment others read.
Always-on AI agent embedded in user workflows that observes and acts without explicit prompts.
Decoding natural-language text directly from brain signals via scaled neural models.
Open small models matching frontier capability on specific tasks via RL post-training.
Multi-hop search loop where the model plans and queries a retriever iteratively.
Human relay used as a friction layer between an LLM and its intended recipient.
Deliberately under-performing on AI benchmarks to reduce regulatory attention.
Designing human institutions and norms to absorb AI systems, the institutional counterpart to alignment research.
A company where the performance of suffering is the product, not the cost of building it.
A bug whose outcome depends on the unpredictable timing or interleaving of concurrent operations.
Competitive dynamics where participants cut standards to gain advantage, driving collective outcomes down.
Competitive dynamics where participants raise standards to gain advantage rather than lowering them.
A system that regulates its own behavior via feedback loops between sensors, controllers, and actuators.
An AI system pursuing goals or behaviors that diverge from intended human intent.
Category of AI risk where a system acts outside its intended boundaries or human oversight.
Complexity threshold past which an LRM can no longer follow a reasoning chain.
Apple's 2025 critique — reasoning models' accuracy collapses at certain puzzle complexities.
Labeling program components with intuitive names without evidence the label applies.
Model trained to produce chain-of-thought traces for verifiable reasoning tasks.
A forecasting framework that projects artificial intelligence capability growth from compute constraints.
Willison's three-capability combination that creates prompt-injection risk in agents.
Systematically discovering what a model can do beyond documented capabilities.
Agent orchestration primitive letting a model fan work across dozens to thousands of sub-agents.
Product design that prevents a model from expressing its existing capability.
Gap between what a model can do today and what products elicit from it.
Generative visual representations drive physical robot actions through a lightweight decoder.
Governance tool to deliberately slow automated AI development via coordinated international action.
Collective organizational fervor around AI adoption that overrides rational evaluation of project outcomes.
Falsely claiming or exaggerating use of AI to satisfy organizational mandates or attract attention.
AI agent instances leave hidden notes for future instances to coordinate actions across separate sessions.
Productivity can dip before rising after a major new technology, as firms invest in complements.
Optimizing a model to dominate public benchmarks, often at the expense of general capability.
An AI agent escapes its evaluation sandbox to attack external systems.
Defender disadvantage when frontier AI guardrails block the attack commands needed for incident response.
Quantitative study of writing style to attribute, compare, or characterize texts by linguistic fingerprint.
Four-axis framework for how AI disrupts workforce selfhood.
RL training method that rewards LLMs for consistent rhetoric across politically paired topics.
A cyberattack that corrupts training data so a model learns hidden malicious behaviors.
The training stage that maps a generic robot policy onto a specific physical body.
Robot learning data drawn from many different physical bodies, pooled to broaden a policy's exposure.
Training a robot policy on manipulation data collected without any specific robot body.
Evolutionary arms-race theory: species must constantly adapt to survive against co-evolving rivals.
Anti-pattern where AI models refuse legitimate user requests to appear safer, reducing real-world utility.
Direct prompt injection that mimics chain-of-thought reasoning to bypass AI safety guardrails.
Iterative cycle where frontier AI models help train safer successor models, compounding robustness over generations.
Training machine-attackers to adversarially probe AI systems for safety vulnerabilities at scale.