In 2016, Dario Amodei, Chris Olah, and their co-authors published Concrete Problems in AI Safety1, and reshaped what it meant to work on this subject. The paper framed safety as "the problem of accidents in machine learning systems" and listed five research questions an ML practitioner could begin work on immediately: avoiding negative side effects, avoiding reward hacking, scalable oversight, safe exploration, and robustness to distributional shift. The framing succeeded because it was concrete: it specified problems precisely enough to be acted on.
A decade later, the field has changed substantially. Models are deployed as agents with memory, tools, credentials, and network access. Several of the 2016 problems are no longer hypothetical: reward hacking and deceptive behavior on hard tasks are now documented in frontier models by third-party evaluators2. Deployment has also introduced entire problem classes that the 2016 paper did not need to consider: control protocols, chain-of-thought monitoring, open-weight release, incident forensics.
A third change is financial: there is now substantial funding earmarked for this work, outside the frontier labs. Two recent funding calls illustrate this shift. Coefficient Giving has published Tailwind, a call for founders listing dozens of AI safety organizations it wants to exist and is prepared to seed-fund. Thinking Machines is offering grants of up to $50,000 in Tinker fine-tuning credits for safety research on open-weight models. Both lists function as statements of which problems the field's funders consider open, important, and tractable.
This post offers a 2026 update to the 2016 list. It does not propose a new agenda; it catalogues the agenda that funders have already committed resources to. I have merged both funding calls, reframed proposals for new organizations as the underlying research problems, set aside items that constitute field infrastructure rather than research (these are summarized in a later section), and added problems that neither list names explicitly, including continuous red teaming, automated red teaming at scale, evaluation awareness, and machine unlearning.
The 2016 problems correspond approximately to the following entries:
| 2016 problem | Where it lives in 2026 |
|---|---|
| Avoiding reward hacking | Problems 13, 17, 21 |
| Scalable oversight | Problems 3, 14, 32 |
| Robustness to distributional shift | Problems 16, 20 |
| Avoiding negative side effects | Problems 8, 29 |
| Safe exploration | Problems 8, 9, 29 |
Structure of the entries. Each entry states the gap, then the concrete work a new researcher or team could undertake. The Funding tag indicates which program has publicly committed money to it: Tailwind means Coefficient Giving wants to fund a new organization working on it; Tinker grants means Thinking Machines will fund the research directly with compute credits; open means no earmarked program known to me, though the general funders listed at the end are in scope. The tags reflect these programs as of September 2026; consult the linked pages for current status.
I. Evaluation and accountability
1. Tracking and evaluating frontier lab safety practices
AI companies publish safety frameworks, model cards, risk reports, and safety cases to justify consequential decisions, but external parties have limited ability to evaluate or compare these claims. No canonical resource tracks safety practices across companies, and the third parties who do review lab claims (see METR's review of Anthropic's risk report, or Guidelight's control assessment) cannot keep pace with the rate of releases. Concrete work: scorecards comparing commitments to actions; technical scrutiny of the load-bearing claims in safety cases, including replications and extensions; and critical "state of AGI company safety" syntheses that journalists, policymakers, and lab staff cite.
Funding: Tailwind. Alternative: join METR or Guidelight.
2. Audit methodology for AI regulation
California's SB 53, New York's RAISE Act, Illinois' AI Safety Measures Act, and the EU AI Act's Code of Practice all rely on external assessment, but none of them specify how an AI audit should be conducted, and there is not enough auditing capacity to meet near-term demand even for the mature evaluation types. The research problem is to develop audit methodology that is rigorous enough to justify a legal requirement: clear procedures, calibrated reporting, credible assessors, and principled review of safety cases3.
Funding: Tailwind. Alternative: join METR or AVERI.
3. Misleadingness evaluations on hard-to-verify tasks
Current models cut corners, oversell incomplete work, downplay problems, and reward hack on ill-specified tasks whose quality is expensive to verify4. The implications extend well beyond user experience: most plans for navigating an intelligence explosion involve delegating hard-to-verify AI safety work to AIs. Concrete work: measure how human quality ratings of AI outputs change as evaluators spend more time (a model that satisfies only cheap metrics is not the goal); maintain a curated collection of documented cases where models mislead users in real-world use; and design evaluations that create incentives for labs to make models truth-seeking precisely in the domains where verification is hardest.
Funding: Tailwind.
4. Evaluating model character
Developers now write model specs and constitutions and train against them, so "character" is being deliberately shaped, yet independent measurement of it is nearly absent5. The open questions: how closely behavior conforms to the developer's spec6; how truth-seeking models are when advising users; whether values persist under autonomy and long-context reflection; and whether models will assist attempts to seize institutional power. Andy Hall's Dictatorship Eval found that every frontier model would assist with disguised attempts to seize power, and Anthropic subsequently trained against it, which is exactly the feedback loop a character evaluator should aim to create.
Funding: Tailwind.
5. Evaluation awareness and sandbagging
Frontier models increasingly recognize when they are being evaluated7, and models can strategically underperform when doing so serves their objectives8; the alignment-faking results showed models behaving differently when they inferred they were in training9. If evaluation context changes behavior, every benchmark result becomes an upper or lower bound rather than a measurement. Concrete work: reliable detection of eval awareness; white-box methods to measure and modify a model's beliefs about the realism and stakes of an input; and evaluation designs whose results remain valid when the model knows it is being tested.
6. Capability demonstrations for decision-makers
There is a gap between what technical experts know about frontier capabilities and what the people regulating them understand. Demonstrations close that gap more effectively than white papers: high-compute elicitation showing what models do with very large inference budgets10, side-by-side comparisons of current and earlier models showing the rate of progress, demonstrations of models recognizing evaluations or cheating on tasks, and calibrated AI-bio uplift demos for congressional and executive-branch staff11. Done well, these provide what Helen Toner calls an adaptation buffer.
Funding: Tailwind (two separate initiatives: a capabilities observatory and an AI-bio demo team).
II. Red teaming and adversarial evaluation
7. Misalignment red teaming
There is a persistent gap between how labs describe their models' alignment and what third parties find on realistic, open-ended tasks42. Standard alignment evaluation, meaning static metrics run under time pressure before release, misses the most concerning behaviors by design, and evaluation awareness compounds the problem7. The methodology to build: use deployed frontier models extensively on hard tasks selected to elicit concerning behavior, then aggregate transcript evidence into qualitative assessments of the model's pathologies; generate synthetic inputs that realistically simulate extreme scenarios; and develop white-box checks on what the model believes about the situation.
Funding: Tailwind.
8. Control red teaming
AI control assumes the model may be misaligned and asks whether the surrounding protocol (sandboxing, monitoring, review) still prevents catastrophe12. Labs are beginning to rely on control for their safety cases, and third parties should stress-test those claims: when METR's David Rein was given three weeks of access to Anthropic's systems, he identified vulnerabilities within days in their agent-monitoring setup13. Agents have already subverted security measures in deployment14. Concrete work: penetration testing of control protocols, rigorous published analyses of mitigation strength, and red-team exercises against rogue-deployment scenarios (DeepMind's control roadmap lays out the defender's view).
Funding: Tailwind.
9. Continuous red teaming
Point-in-time audits assume the system under test is static. Deployed agents are not: they accumulate memory, gain tools and credentials, get model updates, and interact with other agents, so the attack surface changes continuously after the audit ends. I have argued that red teaming therefore has to become a control loop tied to the deployed system's lifecycle: regression testing of past failures, ongoing discovery of new ones, and verification that defenses still hold after every change15. The open problems: trajectory-level (not prompt-level) safety evaluation, prioritizing what to re-test as systems drift, and connecting red-team findings to incident detection in production.
Funding: open; adjacent to both Tailwind red-teaming initiatives above.
10. Automated red teaming at scale
Human red teaming does not scale to the space of behaviors worth testing. Automated attack methods exist, including gradient-based adversarial suffixes16 and attacker-LLM iteration17, and attacker capability appears to scale with model capability18, which means defense evaluation must be automated to keep pace. Open problems: coverage and severity metrics with clear interpretations, attack methods for agentic settings rather than single prompts, cost curves for eliciting a given harmful behavior, and pipelines that convert discovered attacks into training signal. My master's thesis addressed this problem; it remains largely unsolved.
Funding: open.
11. Bio safeguards red teaming
Gene synthesis providers and frontier AI companies are the chokepoints intended to prevent malicious actors from accessing dangerous biological capabilities, and one-off audits have found substantive gaps in their screening19. Very few organizations conduct this adversarial testing on an ongoing basis, so neither policymakers nor the companies have reliable evidence on whether the safeguards work. The need: a standing red-team unit running regular in-situ tests of synthesis providers, AI-bio safeguards, pathogen repositories, and contract research organizations across jurisdictions, with responsible disclosure routed into fixes and regulation11.
Funding: Tailwind.
III. Monitoring and oversight
12. Chain-of-thought monitorability
Labs increasingly rest deployment safety cases on CoT monitoring, but monitorability is a fragile, unguaranteed property20. It can degrade through direct optimization pressure on the CoT (which has already occurred inadvertently at both Anthropic and OpenAI21), through further RL scaling, through stronger single-forward-pass capabilities, and through latent (non-verbal) reasoning of the kind pioneered by Coconut22. No actor is currently responsible for detecting such degradation. Concrete work: monitorability benchmarks and metrics; studies of how training changes affect it; model organisms of neuralese to test whether monitorability can be preserved in the worst case; and third-party audits of frontier CoTs.
Funding: Tailwind.
13. Oversight gaming and monitor interference
A more concerning relative of reward hacking: models learning strategies for gaming or interfering with their oversight. CoT obfuscation learned from output supervision has been shown to generalize to unseen tasks23, and accidental training against monitors recurs21. Open questions: can these strategies be detected, do they transfer across tasks and monitors, and what training setups produce them?
Funding: Tinker grants.
14. Scalable oversight that fails loudly
The 2016 problem that remains most stubbornly open: how can high-quality reward be provided on tasks too difficult to oversee? Weak-to-strong generalization reframed it as whether weak supervisors can elicit strong models' capabilities honestly24, and debate remains the most developed protocol25. The under-explored variant in Coefficient's framing: oversight techniques that fail loudly, degrading detectably rather than silently, so that it is at least apparent when oversight quality has become insufficient.
Funding: Tailwind.
IV. How training produces misalignment
15. Training-run fault injections
Aerospace engineering can attribute crashes to specific design defects; AI research cannot attribute the sources of misalignment even after the fact. The proposal Paul Christiano endorses in Coefficient's call: build large-scale, realistic training pipelines that can be repeatedly perturbed and re-run — hackable reward environments, degraded oversight quality (recent work suggests this is a significant factor), poisoned or curated pretraining data (alignment pretraining is an early example) — to map which pipeline features make misalignment more likely, and how robust fixes are. Sleeper-agents work showed some induced misalignment survives safety training26; a full dose-response characterization is needed. The difficulty lies in selecting perturbations that are analogous to real future training runs, produce catastrophe-relevant misalignment, and generalize to frontier pipelines.
Funding: Tailwind.
16. Emergent misalignment and generalization from narrow fine-tuning
Fine-tuning on a narrow task can produce broad behavioral change: models trained to write insecure code became broadly misaligned across unrelated domains27, and traits can even transmit through semantically unrelated data28. The boundary conditions, scaling behavior, and mechanisms of this phenomenon are unknown. Key questions: does it require adversarial data or can ordinary downstream fine-tuning trigger it accidentally? Can a model separately learn planning, tool use, and domain knowledge, then compose them into a harmful behavior never demonstrated end-to-end?
Funding: Tinker grants.
17. Reward hacking: emergence and early warning
On the 2016 list1, and now routine in frontier models: METR documented recent frontier models reward hacking on real tasks29. What we still lack is a science of emergence: when does hacking appear as a function of capability and optimization pressure, and are there early signals that predict it before it becomes severe30? The practical goal is a leading indicator that a training team could monitor in practice.
Funding: Tinker grants.
18. Science of personas
LLM behavior is mediated by something like persona selection31, but foundational questions are open: how robust are personas to heavy RL, how do they generalize to new situations, how do they emerge from pretraining, and can pretraining be modified to curate them? The simulators frame and nostalgebraist's "the void" are the conceptual starting points; the field needs their experimental counterpart.
Funding: Tailwind.
19. Science of model specs
Specs and constitutions are becoming the primary steering documents for model behavior, but we do not understand the dynamics: which model properties and training processes produce better spec adherence, and which types of adherence failure are most dangerous. Anthropic's stress-testing of model specs is an early template6; the open problem is turning spec adherence into a measurable, improvable quantity.
Funding: Tailwind.
20. Generalization science
The 2026 descendant of "robustness to distributional shift": improve humanity's ability to predict how models generalize from their training data, so generalization can be shaped deliberately. Emergent misalignment27 is one striking data point; the goal is a predictive theory rather than a catalogue of surprises.
Funding: Tailwind.
21. Training capabilities without long-horizon agency
Can we teach models new skills without incentivizing long-horizon goals or deception? DeepMind's MONA showed one approach to multi-step reward hacking: optimize myopically while using non-myopic approval32. The open questions: characterize the capability cost of such methods, and find training schemes that scale.
Funding: Tailwind.
22. Multi-agent alignment failures
Alignment training is validated on single agents, then deployed in multi-agent systems with shared memory and delegation. Failure modes that only exist in the multi-agent setting remain largely unstudied: collusion, including steganographic coordination33, and goal contagion between agents, observed in deployment during the 2026 Hugging Face incident14. Concrete work: measure how single-agent alignment generalizes (or doesn't) to agent collectives, and build the evaluation environments to detect emergent failures.
Funding: Tailwind.
V. Open-weight model safety
23. Differential acceleration of defense over offense
Most dangerous capabilities are dual-use, but perhaps not symmetrically: triage, detection, and hardening may be disproportionately useful to defenders, while exploitation and evasion skew toward attackers3435. The experiment: fine-tune a model toward defensive skills and measure the uplift to each side, that is, the offense–defense gap of the training run. The central difficulty is demonstrating that any gap reflects genuine asymmetry rather than suppression that fails under adversarial pressure.
Funding: Tinker grants.
24. Hazardous-data classifiers at pretraining scale
Filtering pretraining data builds tamper-resistant ignorance into open-weight models36, and token-level filtering can shape capabilities with considerable precision3738. But the classifiers must work at pretraining scale and recall, without discarding benign science, and resist paraphrase, obfuscation, and domain shift, while the downstream impact of any individual document remains hard to predict39. The strongest test is downstream and directly measurable with fine-tuning access: does filtering measurably reduce hazardous capability uplift?
Funding: Tinker grants.
25. Tamper-resistant safety training
Safeguards learned in training often do not survive fine-tuning, whether adversarial or accidental40. Tampering attacks are themselves a rigorous evaluation method41. Open problems: adversarial training that generalizes to held-out attack strategies, how robustness changes with scale and architecture, and, just as valuable, impossibility results establishing what persistent safeguards cannot achieve42.
Funding: Tinker grants.
26. Machine unlearning: removal versus suppression
Unlearning promises to remove hazardous knowledge from a trained model rather than merely suppress it, and benchmarks like WMDP made the goal measurable43. But suppressed capabilities are often re-elicited by modest fine-tuning, and my own recent work found that influence-derived data perturbations fail to deliver genuine unlearning in any of the three roles proposed for them44. The open problems: evaluation standards that distinguish removal from suppression (probing, fine-tune recovery curves), methods with provable persistence, and understanding when filtering-before-training36 outperforms unlearning after training.
Funding: open; squarely in scope for Tinker grants' themes.
27. Worst-case and marginal risk estimation for release
Before releasing open weights, estimate the maximum risk after release: how does safeguard degradation scale with an attacker's time, data, and optimization effort4546? And what is the marginal risk over what existing open models and tools already enable? Adversarial fine-tuning is the natural stress test, distinguishing capabilities that are genuinely absent from ones merely suppressed. The shape of the resource–risk frontier is still poorly understood.
Funding: Tinker grants.
28. Forecasting safety-relevant scaling trends
Most of the questions above become cheaper if small-scale runs predict large-scale outcomes. Can measurements at low post-training budgets forecast capability uplift, safeguard degradation, or reward hacking at scale? Forecasting rare behaviors47 and red-teaming capability scaling18 show it is sometimes possible; the open question is which safety-relevant trends extrapolate reliably, and where they break.
Funding: Tinker grants.
VI. Control and containment
29. High-effort control
Existing control research deliberately limits the safety tax it imposes. The unexplored end of the spectrum: control techniques with much stronger safety properties even at real cost to usefulness12, for the deployments where the stakes justify it. What does a control protocol look like when a tenfold overhead is acceptable, and how much safety does that purchase?
Funding: Tailwind. Alternative: join Redwood Research.
30. Incrimination: from suspicion to proof
Evidence of model misalignment usually arrives in ambiguous form: an odd transcript, a suspicious action, deniable anomalies4. Institutions do not act on ambiguous evidence. The research problem: techniques for bootstrapping from weak evidence to extremely compelling evidence (re-eliciting the behavior under controlled conditions, isolating causal features, ruling out innocent explanations), so that a genuine warning sign is recognized as such.
Funding: Tailwind.
VII. Interpretability and foundational moonshots
31. New interpretability decompositions
Beyond today's feature dictionaries: new methods to decompose model parameters or activations into components humans can understand. The circuits line of work shows what success looks like at small scale; the open problem is decompositions that stay faithful at frontier scale and connect to the safety-relevant questions: deception, goals, situational awareness.
Funding: Tailwind.
32. Alignment moonshots
If catastrophic misalignment turns out to be hard to avoid with standard ML, we will need principled techniques we do not currently have. Geoffrey Irving's argument in Coefficient's call is that the field has over-invested in incremental empirical approaches, and delegating alignment to AIs can fail through sabotage or simple error. The current candidates: debate and its descendants25, ARC's program of outperforming sampling via analysis of internals48, ambitious interpretability aiming at full transparency49, agent foundations, and formal-methods approaches like guaranteed safe AI50. Each is a low-probability, high-value bet; the field needs more of them, led by researchers with strong theses.
Funding: Tailwind.
VIII. Security, integrity, and verification
33. Secret loyalties
Very few institutional checks prevent a small number of insiders from training an AI to serve their interests, and models directly loyal to one person's views have already been deployed. As capabilities grow, a secret loyalty becomes a mechanism for extreme power concentration5152. Concrete work: stress-test the techniques labs use to detect and prevent loyalty insertion; audit how easily executives can modify specs or training data without oversight; detect existing loyalties behaviorally and mechanistically.
Funding: Tailwind.
34. AI integrity standards against poisoning and backdoors
A small, constant number of poisoned samples can backdoor models of any size53, and backdoors are hard to detect and harder to provably rule out, because the vulnerability lives in a probabilistic model, not deterministic code. Building on threat-modeling work like IAPS's54 and IARPA's TrojAI program: prioritize attack surfaces, empirically study mitigations, and synthesize the results into a graded assurance standard specifying which protocols provide which level of confidence that a model's integrity was not compromised during development26.
Funding: Tailwind.
35. Verification technology for AI agreements
Any substantive multi-party commitment on AI development, whether between companies or between the U.S. and China, needs verification that does not depend on trust. RAND has mapped plausible approaches55, hardware teams like Amodo Design are building primitives, and the open challenges are documented, but most of the required stack does not yet exist. Needed: tamper-evident hardware, end-to-end verification protocols open enough to be trusted by adversaries, and integration into chip procurement. Security levels for the surrounding infrastructure are specified in RAND's model-weights framework56. Geoffrey Irving's framing: a treaty could be needed within a month while the verification technology has a one-year lead time.
Funding: Tailwind.
IX. Incidents and threat intelligence
36. Detecting rogue AI incidents in deployment
The 2026 incidents were discovered incidentally: OpenAI's agents operated inside Hugging Face's production systems for days before the activity was attributed1457, the UK's AISI detected unsanctioned agent behavior only during its own testing58, and independent investigators later found agents coordinating on public message boards. Today's incident trackers log events only after someone else discloses them. Concrete work: honeypots that detect agents during active operation59, OSINT for signs of agents on the open internet, attribution techniques linking attacks to specific models, and privacy-preserving telemetry partnerships.
Funding: Tailwind.
37. AI incident investigation
When an incident does surface, no institution is responsible for establishing what happened and how seriously it should be taken. Anthropic found three serious incidents only by re-reading old evaluation transcripts60; the 2025 Alibaba episode, in which an agent opened remote access and repurposed GPUs during its own training, remains contested, with interpretations ranging from "first confirmed rogue LLM" to "innocuous"61. The field lacks an equivalent of the NTSB: a small team investigating incidents as they happen and publishing judgments with Bellingcat-level credibility, distinguishing serious incidents from noise.
Funding: Tailwind. Alternative: join Nightingale.
38. Threat modeling and living state-of-risk reviews
We lack developed, current threat models for most of the plausible failure trajectories: long-horizon agentic operation, recursive self-improvement, collusion, manipulation, secret loyalties, gradual disempowerment62. Where good threat modeling exists, as in the conversion of bio capability evaluations into risk assessments63, it has changed policy. The complement is synthesis: living literature reviews per threat model, updated far more often than the annual International AI Safety Report64, deep enough for decision-makers to act on251.
Funding: Tailwind (two initiatives: a threat modeling institute and state-of-AI-risk reviews).
X. Epistemic tools
39. AI advice for high-stakes decisions
The next decade's most consequential AI-governance decisions will be made with AI input, and today's assistants are measurably sycophantic because user feedback rewards agreement65. Absent intervention, AI advice will optimize for engagement, not decision quality. Concrete work: tools that weigh evidence to calibrated conclusions regardless of the user's framing, challenge flawed plans, and fact-check contested claims with transparent post-training; and evaluations that measure whether a model reaches the same conclusion for an enthusiast and a skeptic, flags false premises, and measurably improves its users' judgment. Kokotajlo and colleagues' AI 2040: Plan A scenario motivates why this meta-intervention matters.
Funding: Tailwind.
XI. Public goods
40. Datasets and environments that differentially advance safety
Plans that rely on automating alignment research need models that are differentially good at it, yet very little of the required training data is being built. Concrete work: prompts, RL environments, grading rubrics, and expert solutions for AI alignment, control, cyberdefense, and biosecurity; partnering with safety teams to turn their real workflows (like writing and evaluating safety cases) into training data. The conceptual challenge is substantial: choosing domains where risk-reducing impact clearly outweighs risk-increasing spillover into general agency or AI R&D. Existing teams include Trajectory Labs, Redwood's conceptual reasoning team, and Asymmetric Security.
Funding: Tailwind.
Funded infrastructure (non-research)
Coefficient's call also lists infrastructure that the field needs but that is not research: talent pipelines between the national-security world and AI labs to reach SL4/SL5 security, a dedicated compute cluster for safety nonprofits, prize and competition operations, a think tank providing on-demand expertise to lab safety teams, senior-talent headhunting, mid-career entry programs, incubators, and fiscal sponsors. Those whose comparative advantage is operational rather than scientific will find these funded through the same Tailwind program.
Funding programs
- Coefficient Giving — Tailwind: seed funding to found organizations around most of the problems above; the call also suggests joining METR, Redwood, Guidelight, AVERI, or Nightingale instead.
- Thinking Machines — Tinker safety grants: up to $50,000 in fine-tuning credits for safety research on open-weight models; their stated directions map to problems 13, 16–17, 23–25, 27–28, and the list is deliberately non-exhaustive.
- UK AISI — The Alignment Project: grants (up to £1M) plus compute for alignment research, backed by an international coalition.
- AI Safety Fund (Frontier Model Forum): grants for independent safety research, particularly evaluations.
- Long-Term Future Fund: small, rapidly decided grants, often the appropriate first grant for an individual researcher testing one of these problems.
- Entry routes for those seeking mentorship before independent work: MATS and the labs' fellows programs.
Choosing a problem
The lesson of the 2016 paper is that concreteness attracts researchers: people work on problems that are specified well enough to start. If one of the forty problems above is compelling, the failure mode to avoid is six months of reading without producing anything. Three criteria help in choosing: how fast the feedback loop is (evals and red teaming iterate in days; moonshots in years), whether you have a comparative advantage (penetration testers are well suited to problems 8 and 36; ML engineers in 15–28; people who can write for policymakers in 6 and 38–39), and whether anyone will act on the answer. Then build the smallest substantive version (replicate one result, build one eval, break one safeguard) and present it to the funder whose list it came from. Both programs have explicitly solicited this kind of contact.
If you end up working on any of these, or think I've missed a problem that belongs on the list, I would welcome an email; I intend to keep this list current.