From Aerospace Manuals to AI Explainer Videos: Andrej Karpathy Outlines the Future of Human-AI Interaction
SAN FRANCISCO — In a wide-ranging post published on X (formerly Twitter) this Thursday, prominent AI researcher and Anthropic pretraining team member Andrej Karpathy detailed a progressive framework for maximizing the utility of large language models (LLMs). By rethinking how artificial intelligence presents information—moving systematically from strictly constrained prose to interactive diagrams, custom HTML pages, and AI-generated video explainers—Karpathy argues that human labor is rapidly shifting away from manual execution and toward high-level oversight and conceptual synthesis.
The core of Karpathy’s proposal centers on a surprising historical artifact: ASD-STE100, a specialized, highly regulated form of English originally designed to make complex aerospace maintenance manuals unambiguous and easy to read. By instructing language models to adopt these strict stylistic parameters, or by pushing outputs into richer visual formats, users can drastically reduce cognitive fatigue and process dense information with unprecedented efficiency.
As AI capabilities scale, Karpathy’s insights offer a compelling glimpse into the evolving workspace, where the bottleneck of human productivity is no longer generating text, but successfully digesting and validating machine intelligence.
Main Facts
The foundational premise of Karpathy’s recent viral post is that modern LLMs possess a deep understanding of human communication formats, yet humans often default to reading standard, unstructured paragraphs—an inefficient medium for complex concepts. To combat this, Karpathy outlines a four-step evolutionary ladder for consuming AI output, declaring each subsequent step "even better" than the last:
- ASD-STE100 (Simplified Technical English): Leveraging strict aerospace documentation standards to eliminate ambiguity, enforce active voice, and severely limit vocabulary choices.
- Structural Diagrams: Bypassing text entirely for certain concepts by asking models to generate clean visual schematics that are easier to parse.
- Interactive HTML Pages: Requesting fully rendered, self-contained web applications instead of static Markdown text, allowing for dynamic exploration of data.
- Custom Explainer Videos: Generating automated, animated videos complete with custom narration—modeled after popular educational channels like 3Blue1Brown—using programmatic tools and text-to-speech APIs.
Karpathy, a founding member of OpenAI who transitioned to Anthropic’s pretraining division in May, suggests that these techniques are harbingers of a broader shift in knowledge work. "As models do more of the legwork," Karpathy wrote, "a lot more of our work will rise up the abstractions into oversight and understanding."
Crucially, Karpathy also encourages users to lean into the creation of "large, custom, discardable software artifacts"—disposable code and interactive environments tailored to a single query that "would have never made sense to create before" due to prohibitive human labor costs.
Chronology of Events
To understand how Karpathy arrived at his current framework, it is helpful to trace the timeline of his recent public commentary and professional moves regarding AI interfaces and tool usage:
- May 2026 (Early Month): Karpathy shares an X post authored by Thariq Shihipar of Anthropic, detailing the "unreasonable effectiveness" of using HTML over Markdown within engineering workflows. This philosophy is subsequently formalized in a Claude Blog post later that month.
- May 19, 2026: Media outlets report that Andrej Karpathy has officially joined Anthropic’s elite pretraining team, marking a major career move following his foundational work at OpenAI and his subsequent independent educational projects.
- Late May to Early June 2026: AI community members experiment heavily with agentic coding environments, browser-based renderers, and custom local software workflows, setting the stage for more advanced multi-modal presentation requests.
- Thursday (Current Week): Karpathy publishes his definitive X thread introducing ASD-STE100 as a prompt-engineering hack, alongside his bullish stance on automated, 3Blue1Brown-style video generation tools utilizing text-to-speech integrations like ElevenLabs. Within hours, open-source developers begin circulating auxiliary GitHub repositories and parsing skills designed to align AI agent text inputs with strict technical writing guidelines.
Supporting Data and Technical Context
Karpathy’s framework relies on a mix of legacy industrial standards and cutting-edge generative software pipelines. Examining the individual components reveals why these methods yield superior cognitive returns for human reviewers.
What is ASD-STE100?
Maintained by the AeroSpace and Defence Industries Association of Europe (ASD), Simplified Technical English (STE)—officially designated as ASD-STE100—is an international standard originally created in the 1980s to help maintenance engineers worldwide understand complex aircraft manuals without language barriers.
Key attributes of STE include:
- Controlled Vocabulary: A restricted dictionary of roughly 900 approved words, where each word typically has only one primary meaning and one part of speech.
- Rigid Structural Rules: Strict limitations on sentence length (e.g., maximum 20 words for procedural sentences), mandatory active voice, and the elimination of auxiliary verbs where possible.
- Clarity Over Style: The complete stripping away of marketing jargon, metaphors, and stylistic flourishes in favor of unvarnished precision.
Because LLMs are trained on vast web corpora that include technical manuals and specifications, they naturally internalize the syntax of STE. Karpathy noted that because the standard is "quite stringent," he occasionally modifies his prompt to request text that is "80% of the way to ASD-STE100," softening the rigidity while retaining the clean writing style.
In the wake of his post, open-source developers quickly highlighted existing community resources, such as GitHub repositories housing STE-filtering skills designed to process text destined for AI agents (though developers note these filters are optimized for technical documentation rather than creative marketing copy).
The Rise of HTML and Dynamic Artifacts
Karpathy’s endorsement of HTML outputs builds on a growing consensus among elite software teams. Standard Markdown text is linear and static. By asking an LLM to wrap its response in a functional HTML page complete with CSS styling and JavaScript interactivity, the reviewer receives a bespoke user interface tailored precisely to the dataset or concept at hand.
Automated Explainer Videos
Perhaps the most forward-looking element of Karpathy’s thread is his enthusiasm for custom video generation. By prompting an LLM to draft a script and layout styled after educational math channels like 3Blue1Brown (known for mathematical animations generated via Python libraries like Manim), and pairing it with an API key from text-to-speech providers like ElevenLabs—or running free local open-source speech models—users can instantly generate multimedia lessons on demand. "This is actually starting to work!" Karpathy marveled.
Official Responses and Industry Reactions
The AI research and engineering communities reacted swiftly to Karpathy’s thread, debating both the practical applications and the philosophical implications of engineering human-readable AI outputs.
While the ASD organization itself has long maintained that no automated tool can fully replace human mastery of STE for safety-critical aerospace documentation, software engineers and prompt architects have embraced the standard as a powerful hedge against "LLM bloat"—the tendency of generative models to pad responses with unnecessary fluff, conversational filler, and vague corporate prose.
Industry analysts point out that Karpathy’s recommendations highlight a fundamental shift in user experience design. For decades, software interfaces forced humans to adapt to the constraints of machines (typing exact commands, navigating rigid menus). Generative AI inverted this by allowing machines to understand natural human language. However, as Karpathy’s post illustrates, human reading comprehension has now become the primary bottleneck in the loop. By forcing the AI to output information in hyper-optimized technical prose, interactive web apps, or animated video, users are essentially designing custom cognitive interfaces on the fly.
Implications for the Future of Work
Karpathy’s observations carry profound implications for the future of knowledge work, software development, and education:
- The Death of Traditional Documentation: As AI models become capable of spinning up interactive HTML dashboards and custom video explainers on demand, static PDF reports and multi-page memos will increasingly look obsolete. Organizations will shift from writing documents to commissioning dynamic data artifacts.
- Upskilling in Oversight: As routine writing, coding, and synthesis are offloaded to algorithms, the value of human labor will concentrate heavily in critical evaluation, conceptual verification, and architectural direction—what Karpathy describes as rising "up the abstractions into oversight and understanding."
- Personalized Education: The convergence of LLMs, programmatic animation engines, and advanced text-to-speech capabilities points toward a future where every student or professional can have complex, multi-modal tutorials generated specifically for their unique learning style and current knowledge baseline.
Ultimately, Andrej Karpathy’s framework demonstrates that getting more out of artificial intelligence is not merely a matter of building smarter models, but of radically reimagining how we consume, visualize, and interact with the intelligence we already have.
