# Next
<div class="pills-container">
<span class="pill">Last Updated: July 30, 2026</span>
</div>
## Why I work on AI safety
I'm quite concerned about what a post-[TAI](https://www.sciencedirect.com/science/article/pii/S0016328721001932) society looks like for the majority of the world. I would've said "post-[AGI](https://www.agidefinition.ai/) society," but I've recently been more and more convinced that [you don't even need AGI to massively disrupt societal institutions](https://www.alignmentforum.org/posts/LpM3EAakwYdS6aRKf/what-multipolar-failure-looks-like-and-robust-agent-agnostic) and cause [astronomical suffering](https://80000hours.org/problem-profiles/s-risks/). More specifically:
1. I think powerful AI (either the existence of a single powerful AI system or the deployment of mildly intelligent AI systems on a massive scale) can trigger [transformative development](https://forum.effectivealtruism.org/topics/transformative-development) in society. This can mean [10x-ing research progress to solve humanity's biggest problems](https://darioamodei.com/essay/machines-of-loving-grace), but it could also mean 10x-ing the process of [disempowerment](https://gradual-disempowerment.ai/), which could lead to large-scale abuses of power. This transformative development can be further accelerated by an [intelligence explosion](https://www.forethought.org/research/three-types-of-intelligence-explosion#summary) (once AI systems gain the capability to improve themselves).[^1]
2. Once the intelligence explosion takes off, [middle powers (and the rest of the world) will be left out by default](https://newsletter.forethought.org/p/how-can-the-middle-powers-avoid-getting) since [power could concentrate to whichever state triggered it, if that state can block technological diffusion](https://www.forethought.org/research/could-one-country-outgrow-the-rest-of-the-world), which may not be achievable. Ideally, we would want to avoid [extreme power concentration](https://80000hours.org/problem-profiles/extreme-power-concentration/) since this exposes us to [risks](https://80000hours.org/problem-profiles/extreme-power-concentration/#why-might-ai-enabled-power-concentration-be-a-pressing-problem) that we (humanity) may not even be prepared to face.
That said, there needs to be active effort in making sure that our institutions are prepared for such transformative development brought about by the intelligence explosion. Working on this problem means both trying to prevent extreme power concentration where that's tractable, and reducing harm to the people who end up living under whatever concentration happens regardless. The former means efforts to [distribute power over AI](https://80000hours.org/problem-profiles/extreme-power-concentration/#there-are-ways-to-reduce-this-risk-but-very-few-are-working-on-them) and [improve decision-making](https://www.forethought.org/research/preparing-for-the-intelligence-explosion#6-agi-preparedness). The latter means [helping middle powers build resilience](https://writing.antonleicht.me/p/how-ai-safety-is-getting-middle-powers) in a world they mostly can't shape. That can mean:
1. Helping middle powers occupy real bottlenecks in the AI supply chain.
2. Making sure that institutions (in general) are capable of [good reasoning](https://www.forethought.org/research/whats-important-in-ai-for-epistemics), since distributing power alone accomplishes little if the institutions holding that power can't reason well.
3. Working on governance interventions to help middle powers secure access and build AI misuse resilience and economic resilience in a world shaped by AI development they can't meaningfully influence. I echo [Anton Leicht's argument](https://writing.antonleicht.me/p/how-ai-safety-is-getting-middle-powers) that safety advocates should pivot away from trying to shape frontier AI from the outside and toward this kind of national strategy work instead (this is tractable and neglected).[^2]
Simply put, I don't want to live in a world defined by extreme power concentration. I don't want to be at risk of [political disempowerment](https://80000hours.org/problem-profiles/extreme-power-concentration/#top) - to not have a say in the decisions that will affect my life and my family's lives and futures. Hence, I want to work to preserve human agency, especially in decisions that matter for our world and could affect the quality of life for people everywhere.
## Bets I think are worth taking
I think the following are bets towards building resilience in a post-TAI/AGI society. In my worldview, that consists of (1) working on distributing power over AI to make sure that the technology does not have a single point of failure, (2) working on improving decision-making so that both AI systems and societies are equipped with the tools to reason well, and (3) working on middle power resilience, since most states will only have this option instead of being able to build leverage.
Note that I could also be wrong about some of the assumptions here, as [[Ethos#Mistakes I've made|I've been wrong many times before]], and I haven't done a comprehensive literature review for all of these. If you want to work on any of these, [reach out on LinkedIn](https://www.linkedin.com/in/llenzl/).
### Working on distributing power over AI
I'm less keen on doing this myself vs the two strategies below considering my own skills and background. However, I do think this is important and someone, even if not me, should work on it.
1. **Making open-weight models tamper-proof and trackable.** [Casper et al.](https://stephencasper.com/open-technical-problems-in-open-weight-ai-model-risk-management/) lay out the open technical problems here, like training data curation, tamper-resistant training, model tampering evals, staged deployment, and provenance. [Casper's living doc](https://docs.google.com/document/d/10XkZpUabt4fEK8BUtd8Jz26-M8ARQ6c5iJCbefaUtQI/edit?usp=sharing) also has concrete project-level ideas that are worth exploring if this space interests you.
2. **Making sure no single actor ends up with too much influence over AI.**
- Designing the terms of contracts between labs and governments to make sure no one actor has too much influence.
- Mandating [transparency](https://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#increasing-transparency) into AI capabilities, how they are being used, model specs, safeguards and risk assessments, so it's easier to spot concerning behaviour.
- Introducing [more robust whistleblower protections](https://www.lawfaremedia.org/article/protecting-ai-whistleblowers) to make it harder for insiders to conspire or for company executives to suppress the concerns of their workforces.
3. **Helping middle powers occupy real bottlenecks in the AI supply chain**. Countries holding a genuine chokepoint like chip lithography, high-bandwidth memory, downstream manufacturing, or biotech production can [force the frontier developer to remain dependent on a plural set of actors rather than becoming self-sufficient](https://asteriskmag.com/issues/15/beware-the-permanent-periphery).
### Working on improving decision-making (and related dynamics)
1. **Improving AI's decision-making, mainly by making sure AI doesn't concede to power-seeking behavior by other (human or AI) agents.**
- [80,000 Hours](https://80000hours.org/problem-profiles/extreme-power-concentration/) already lists the core technical mitigations here well, like red-teaming model specs against power grabs, auditing models for secret loyalties, and [training AI to follow the law](https://law-ai.org/law-following-ai/).
- Making sure that AI supervising other AI goes well and scales well, [whether similar AIs can oversee each other,](https://arxiv.org/abs/2502.04313) and what [cooperative risks](https://arxiv.org/abs/2211.14468) come up if they do.
- Working on making "good" values stable or [self-healing](https://bengoertzel.substack.com/p/goals-that-grow-back) - something like tamper-proof values.
2. **Improving institutional decision-making by increasing understanding and awareness of AI across journalists, communicators, and the public sector.**
- Making sure evaluations used by AI labs and governments have [ecological](https://en.wikipedia.org/wiki/Ecological_validity) and [construct validity](https://en.wikipedia.org/wiki/Construct_validity) by working on more interp-based evals, making sure open-source projects in AI safety like Inspect are audited, and making transcript analysis easier to do.
- Understanding where disempowerment happens right now. Most of what we currently measure about AI's effect on society is either what AI can do (capability benchmarks) or how much it's used (adoption statistics). Neither of which tells us whether humans have actually lost the ability to act without AI. Given the risk of [gradual disempowerment](https://gradual-disempowerment.ai/), it seems natural to build baselines so we are able to track disempowerment longitudinally and act accordingly. [Chooi, Lee, and Li (2026)](https://openreview.net/forum?id=JfMaFoC2FH) propose 6 metrics to measure gradual disempowerment. I also think we should think of other metrics that we can track over time.
- Understanding where disempowerment can happen in the future. Threat models and scenario-mapping are especially helpful in communicating the urgency of why we need to build resilience for a post-TAI/AGI society. Off the top of my head, I think documents like [AI 2027](https://ai-2027.com/), [AI 2040](https://ai-2040.com/), or [Europe 2031](https://europe2031.ai/) are very helpful for middle powers who need to be mobilized toward interventions that prevent AI harms from manifesting.
- Building the talent pipeline for AI diplomacy, since [current safety-aligned mentorship programs funnel talented people who want to work on their home countries' AI strategy into US-focused development work instead](https://writing.antonleicht.me/p/how-ai-safety-is-getting-middle-powers); redirecting that pipeline is tractable and neglected.
### Working on middle power resilience
If extreme power concentration happens anyway, the people living downstream still need to avoid ending up destitute or defenseless against misuse.
1. **Securing leverage-free access to frontier AI** by [offering sites, funding, and electricity in exchange for guaranteed frontier access](https://asteriskmag.com/issues/15/beware-the-permanent-periphery).
2. **Building AI misuse resilience and economic resilience domestically.**
- [Defensive-focused innovation](https://writing.antonleicht.me/p/how-ai-safety-is-getting-middle-powers) against AI-enabled misuse ([def/acc]() in the general case), since middle power governments are less likely to get this right by default and criminals will target middle power populations regardless of what happens at the frontier.
- Fixing the economic trajectory before it becomes destitution. Leicht's point is that [great powers have policy backstops against AI-driven disempowerment (taxation, redistribution) that middle powers mostly don't](https://writing.antonleicht.me/p/how-ai-safety-is-getting-middle-powers), so this needs addressing before growth divergence becomes what he calls existential rather than merely uncomfortable.
- [Flexicurity-style labor protections](https://asteriskmag.com/issues/15/beware-the-permanent-periphery), high-but-temporary unemployment insurance, and generous training programs so workers can pivot without being crushed by the transition.
---
[CV](https://bit.ly/ld_cv) · [Substack](https://halfbakedtheories.substack.com/) · [LinkedIn](https://www.linkedin.com/in/llenzl/) · [GitHub](https://github.com/ramennaut) · [Feedback](https://www.admonymous.co/lenz)
[^1]: In June 2026, Anthropic released a [statement](https://www.anthropic.com/institute/recursive-self-improvement) to call for an "option to slow or temporarily pause frontier AI development to enable societal structures and alignment research to keep up with the advance of the technology."
[^2]: AI harms (including gradual disempowerment) might manifest earlier in middle powers due to weaker governments, and so getting AI wrong in middle power governance can lead to a destabilized world. Leicht argues the path back to any real influence on AI development runs through pursuing national strength for its own sake.