# Aria Han Aria Han is an AI consultant in Los Angeles who builds and repairs AI products for founders and independent builders, and creates internal AI workflows for operations teams. Canonical site: https://ariaxhan.com Projects: https://ariaxhan.com/api/projects.json Writing: https://ariaxhan.com/api/writing.json Full context: https://ariaxhan.com/llms-full.txt # AI for work, AI for humans. Workflows, systems, and everything else to go from “we need AI” to “we have AI.” Aria Han is an AI consultant in Los Angeles who builds and repairs AI products for founders and independent builders, and creates internal AI workflows for operations teams. ## What I build - **Practical AI workflows:** I teach and build AI workflows for founders and independent builders: research, writing, operations, and decision-making systems they can keep using themselves. - **Internal operations workflows:** For companies, I build internal workflows that connect AI to the tools, files, and knowledge already in use. The work starts with the operation as it runs now, including where context gets lost and effort gets repeated. - **AI products for founders:** I work with founders and independent builders from an early idea through a working AI product. The implementation stays close to the person making the decisions, so the product can change as the problem becomes clearer. - **Agentic system architecture:** I built multi-agent coordination when it was still in its infancy: a single prompt fanning out into a family of agents in a coordinated dependency graph, each with its own role and tools, communicating through handoff notes instead of expensive cross-talk. I design structure so agent work becomes artifacts instead of fog. - **Evals, monitoring, and quality layers:** Checks that tell you whether the AI is doing the thing before a customer, teammate, or future version of you finds out the hard way. - **Claude Code / AI coding workflow hardening:** I spend a vast majority of my time talking to Claude Code. I can make the workflow calmer, more accountable, and less like a very expensive chaos machine, including the part nobody warns you about: agents behave very differently against years of legacy code than a project built from scratch. - **Memory, context, and knowledge systems:** Context is, in fact, everything, and it should not depend on whoever happens to remember it that week. I build the memory and knowledge layers that let people and agents hold their context: structured artifacts, richer recall, and transparency into exactly what the AI is referencing. - **AI product review and repair:** If an AI-assisted build mostly works but has become hard to debug, extend, or trust, I can review the product, trace where it is failing, and help turn the prototype into something you can keep building. ## Selected work ## ModelMind AI education felt backwards, so I built the course I wanted people to have first. Most AI courses start with transformers. That is not what most people need first. They need to understand what is happening well enough to have a better conversation with an LLM, write a better prompt, or recover when the model fails. ModelMind is my answer to the question of how to learn AI. It was inspired by Duolingo, which I have kept a streak on for hundreds of days and counting. The app teaches the concepts behind LLMs through daily, gamified exercises. The knowledge comes from thousands of hours spent talking to models as they evolved, plus the research papers and expert writing I kept referencing in my own process. It is completely free. No paywalls, no ads. My only ask is that it helps demystify these models and build mental models that will matter more and more. Proof: Live on the App Store for iPhone, iPad, and Mac. 495 commits of solo development. Stack: React Native · TypeScript · MMKV ## Paper Rooms Safari tabs ate my research life, so I built a library. I have a daily automation that pulls new AI, machine learning, and wildcard research papers into a digest. After months of reading that email, squinting at PDFs, and losing links in Safari history, I got tired of the mess. Paper Rooms pulls research papers in from a link, reformats them into something readable, and organizes them into a library inspired by the way real libraries catalog books. It was built for one very specific purpose: reading research papers without losing them, hating the PDF, or turning a good rabbit hole into browser archaeology. It is also free, with no ads. Proof: Live on the App Store for iPhone, iPad, and Mac. Built and shipped solo in under a week. Stack: Capacitor · Local storage ## our4cuts An iPad and a browser are all you need for a photo booth. Event photo booths are hardware rentals, but a booth is really just a camera, a layout, and a shared gallery, all things a phone browser already has. Scan a QR code, guests shoot four frames in the browser, every strip lands in a live gallery. Weddings, pop-ups, restaurant photo zones. Proof: Live and used at real events. 435 commits of production hardening. Stack: Astro · Cloudflare ## Civic Forges I wanted a city builder where running the city and living a life are the same game. City builders usually stop at zoning and budgets. The person making those decisions disappears behind the interface. A playable Three.js city builder crossed with a city-leader life simulator. Draw roads, zone land, grow a tax base, manage pollution and districts, then advance the mayor's life through choices that change both the person and the city. Proof: Playable in the browser with persistent saves, responsive controls, and 36 automated tests. Stack: Three.js · TypeScript · Cloudflare Workers ## Not Recommended YouTube should begin with a question, not an infinite feed. Recommendation feeds turn an intention into passive scrolling before you have decided what you came to learn or watch. A Chrome extension that replaces YouTube's home feed with one search box, three random Wikipedia detours, and your own saved questions and videos. It also hides Shorts, comments, and recommendation sidebars by default. Proof: Version 0.1.0 is published on the Chrome Web Store. No accounts, analytics, ads, or remote executable code. Stack: JavaScript · Chrome Manifest V3 ## Hearth The notch could be a quiet home for the context I keep reaching for. Open conversations, temporary files, system state, and small obligations live in different places and disappear the moment attention moves. A local macOS companion that opens from the notch into windows, conversation memory cards, a file shelf, system stats, calendar, clipboard history, and local Ollama chat with optional screen vision. Proof: Runs locally on macOS 14+, with deterministic tests for transcript parsing and conversation status. Stack: Swift · SwiftUI · AppKit · Ollama ## HeyContext Context kept disappearing, so we built a workspace around memory. At the time, AI usage was still conversational, not agentic. Going back and forth with a chat assistant was slow and tedious, with context getting lost in the noise. A single user prompt generated a family of agents in a coordinated dependency graph. Each agent had a role, tools, and structured artifacts to work on. They communicated through A2A notes, so agent D could see what agents A, B, and C had learned without paying the time and token cost of direct cross-agent conversation. My favorite system was the crystal dam: conversational context accumulated until it hit a token count or time threshold. When the dam broke, we processed it into stardust, shards, and crystals, memory artifacts users could actually see and inspect. Proof: Went live with hundreds of users within a month, no ad spend. Stack: FastAPI · Redis · Convex · Agno · Next.js ## HeyContent Creators had context everywhere and nowhere, so we tried to bring it into one place. A creator's Instagram, YouTube, Gmail, and notes don't know about each other, so no tool could answer a question about the whole body of work. It started with a hackathon project called Content Creator Connector and became a realization that context is, in fact, everything. Powered by plenty of Monster Energy drinks, pure conviction, and a lot of Cursor sessions, we built a platform that integrated with YouTube, Instagram, and Gmail. My favorite part was the conversational onboarding. It asked targeted, adaptive questions, generated a customized persona, then let the user see and edit it. That persona became the context layer for scripts, posts, and ideas that sounded like something the creator would actually write. Proof: 5+ platforms integrated with real-time sync; the memory layer survived into the next company. Stack: Embeddings · Semantic links · Real-time sync ## Brink Mind Mental health tools ignored the body, so Brink Mind brought heart data into the room. Mental health apps ignore the body. Heart rate and HRV carry signal a journal never captures. Brink Mind linked to the Apple Watch, using biometrics and journal entries to provide safer, more grounded support. It was the end of 2024, still early enough that I was teaching myself SwiftUI, UI/UX, product design, and how to be a CEO at the same time. Proof: Reached TestFlight with working voice, biometrics, and on-device inference. Stack: Swift · Python · Core ML · HealthKit ## site-spec A website can look finished while everything machines need is quietly broken. The browser only shows the visible layer. Search crawlers and AI answer engines depend on another one: robots policy, structured data, response headers, accessibility semantics, sitemaps, and real server-rendered content. site-spec audits that invisible layer. It crawls a live URL or a local build, reads the pages and HTTP headers machines actually receive, and reports concrete errors across AI searchability, SEO, structured data, accessibility, privacy, security, performance, and link integrity. It also includes a deterministic site compiler. A validated SiteSpec becomes deployable HTML with the machine-readable foundation built in, keeping model-written markup and invented facts out of the compile path. Proof: v0.2.0 ships live-URL and local-directory audits, JSON output, CI-ready exit codes, and deterministic site builds. Stack: TypeScript · Node.js · Vitest ## Nexus Office Agents that keep working without you need a place where you can see and answer them. Scheduled agents, issue pipelines, and permission gates are invisible by default, so work can stop silently in a terminal nobody is watching. A native Mac and phone interface for live agent threads, repositories, issues, pull requests, scheduled flows, and permission gates. Every consequential button acts on the real system and every gate answer carries the exact question ID. Proof: Public MIT repository with fixture-driven UI, deterministic runtime state, and live Mac and phone clients. Stack: SwiftUI · Python · GitHub CLI · Tailscale ## renderstate A screen should be inspectable in every meaningful state before the whole app can run. React screens hide their important states behind routers, backends, accounts, and setup that make visual review slow and incomplete. A deterministic UI-state renderer that discovers React screen components, reads their TypeScript prop contracts, derives strong state combinations, renders each in isolation, and emits screenshots plus a neutral manifest. Proof: Version 0.1.2 is published on npm with multi-viewport capture and a static board viewer. Stack: TypeScript · React · Vite · Playwright ## KERNEL Claude Code kept starting over, so I built memory, rules, and receipts around it. Every session starts from zero and every best practice is folklore. Agents need persistent memory and rules that prove themselves. KERNEL gives Claude memory, deterministic hooks, skills, and a way to prove which workflows actually work. Specialized agents, SQLite-backed workflows, validation gates. Installs through Claude's plugin marketplace, mirrors into Cursor and Codex. Proof: 360 commits since January 2026. Distributed through the plugin marketplace, used daily in my own consulting work. Stack: Claude Code · SQLite · Shell ## llm-bench Leaderboards were not answering my questions, so I made tests that did. Leaderboards don't answer the only question that matters: will this model hold up on your actual work? Real workflow tasks: extraction, code, planted bugs, email drafting, prompt injection, each graded by a programmatic verifier. Works with Ollama, Apple Intelligence, Claude CLI, Bedrock, any OpenAI-compatible endpoint. Proof: 21 tests, programmatic verifiers, published model comparisons including Opus 4.8 vs 4.7 vs Sonnet vs Haiku. 148 commits. Stack: Python · Ollama · Bedrock · Claude CLI ## the-agent-library Prompt collections kept rotting, so I turned repeated workflows into portable skills. Prompt collections rot. The useful unit is a workflow with a trigger, steps, and a definition of done that any agent can load. A curated set of portable skills for getting real work out of AI agents, built for Claude, Codex, and any agent that can load a skill file. Most of it isn't code-specific: checking your own work, planning, brainstorming, research, writing, shipping. Each skill is a standalone workflow with a clear trigger and a SKILL.md. Real patterns that survived months of usage, constantly updated. Proof: 39 skills, each extracted from repeated real-world use, MIT licensed. Stack: Claude · Codex · Agent Skills ## model-familiarity-engine I wanted to know what a model had earned across a working relationship, not where it ranked. Single-shot benchmarks don't capture how a model behaves across a real working relationship. Onboards language models by simulating real user conversations, then builds evidence-backed model cards from observations instead of a ranking. All benchmarks are drawn from real conversation transcripts. The replay-bootstrap loop is shipped: known-outcome tasks, redaction, replay, model cards built from what was actually observed. Proof: Replay-bootstrap loop shipped: redaction, replay, observed model cards. MIT licensed. Stack: Python · Bedrock · Ollama · Claude CLI ## metabrain Memory gets noisy unless it has to prove itself. Most memory tools remember. Almost none of them learn. Storage without a promotion gate becomes noise. A zero-dependency SQLite layer that closes the loop: patterns graduate into hypotheses, outcomes test them, and only what holds up becomes preference. Proof: Published: pip install metabrain. Zero dependencies by design. Stack: Python · SQLite · Zero-dependency ## agentmailkit A scheduled email should look the same every day, even when a model writes the words. Cloud assistants can schedule an email but cannot read the files on your laptop or send from your own inbox. Local agents can do both, and then drift: the same job returns a different shape every morning, so you stop trusting it and stop reading it. An email is two files: a JSON job and a markdown prompt. Everything type-specific is a named plugin, so adding a digest means adding data, never code. The split that makes it stable is that the model writes only the words. A deterministic renderer owns every piece of presentation, so the same job produces the same shaped email on every run and only the sentences change. A seen-ledger keyed by job and item URL filters sources before the prompt is ever built and records only after a send succeeds, so day two never repeats day one. Proof: Five example jobs ship with it, so the first command after installing renders real digests from live weather, news, and arXiv feeds. 36 tests, offline and deterministic. Dry runs and the quickstart gallery cannot send, enforced in code rather than by convention. Stack: Python · Append-only JSONL ledger · Zero required dependencies ## Substrate I wanted to see what an unattended creative pipeline would become if it kept going. What does an autonomous creative pipeline actually produce over a year? Almost nobody runs the experiment long enough to find out. A generative gallery where Claude Code agents create abstract, interactive computational art through a fully automated daily workflow. Each piece is a single HTML file, roughly 2KB. Proof: 425 pieces and counting, generated daily without supervision, all public. Stack: HTML · CSS · JavaScript · Cloudflare Pages ## latent-diagnostics Correct answers were not enough. I wanted to know whether the model computed something real. Grading answers tells you whether a model was right, not whether it computed something real. Measures attribution graph geometry instead of only grading answers. Task domains show real signatures after controlling for length. Hallucination detection did not survive the same test. The repo keeps the negative results in. Proof: Grammar influence d=1.08 after length control. 108 commits. The failed hypothesis is documented next to the confirmed one. Stack: Python · SAEs · Attribution graphs # Hi, I'm Aria I spend a vast majority of my time talking to Claude Code, reading books, and writing everything from prompts to poetry. I started as a language person. Journalism, essays, stories, poems, research rabbit holes, and constant reading. Computer science did not feel like leaving that behind. It felt like picking up one more language. Then language models arrived, and the language part and the machine part stopped feeling separate. Since then, my career has been a sequence of frictions. I live with something until it annoys me enough to build a system around it, usually one that does the mechanical work so the person does not have to. I don't think my differentiator is memory, evals, or agents. It is why I keep building them. After my third startup, I spent months interviewing for AI engineer roles and doing over a dozen technicals. It accidentally became a tour through the industry's confusion: every company had a different idea of what AI work was supposed to be. What stayed with me was the range of problems. Since then I have worked inside other people's systems as well as my own, with founders, engineers, and teams trying to make AI useful without letting it flatten the work around it. My work covers AI product building and repair, personal AI workflows for founders and independent builders, and internal operations workflows for companies. I follow current models, tools, methods, and research, then review existing work when they make a materially better approach practical. ## What I'm exploring now - **KERNEL:** My Claude Code plugin. Persistent memory, agents that split the work instead of stepping on each other, an experiment engine that proves which workflows hold up. Active, open source, installable. - **llm-bench:** Practical workflow benchmarks for local and API-hosted language models, graded by programmatic verifiers. - **model-familiarity-engine:** Evidence-backed model cards from replayed known-outcome tasks and observed model behavior. - **the-agent-library:** A curated library of portable skills for checking your own work, planning, generating novel ideas, research, writing, work management, and code engineering. # Writing ## [Grok Bot: Cursor’s Bet on Multi-Agent AI](https://medium.com/@ariaxhan/grok-bot-cursors-bet-on-multi-agent-ai-77d69ba8b290) Of all the AI tools I’ve tried, Grok Bot may be the simplest iteration of a “personal agentic team.” The problem is that it keeps breaking. 4 min ## [Your AI Harness Is the Real Product](https://medium.com/@ariaxhan/your-ai-harness-is-the-real-product-f0fabb3614c4) The model writes the code. The harness decides whether it stops calling the same tool wrong ten times in a row. 9 min ## [How to Secure API Keys for AI Agents](https://medium.com/@ariaxhan/how-to-secure-api-keys-for-ai-agents-ca773a66bd84) When the AI asks you for a key, that's the exact moment to stop. The most dangerous habit in AI coding, and what to do instead. 12 min ## [The Agent-Ready Web: A Working Guide to Cloudflare's New Score](https://medium.com/@ariaxhan/the-agent-ready-web-a-working-guide-to-cloudflares-new-score-1ed0fce8d760) I pointed Cloudflare's new agent-readiness scanner at my own site. Zero of thirteen. 12 min ## [I Put ChatGPT in Charge of Claude Code](https://medium.com/@ariaxhan/i-put-chatgpt-in-charge-of-claude-code-7b9bf5bb8ea9) What happens when you use one model to orchestrate another? 5 min ## [Stop Writing Markdown. Start Writing Memory.](https://medium.com/@ariaxhan/stop-writing-markdown-start-writing-memory-e4a69c57caa9) Markdown is optimized for human eyes. Terrible for knowledge agents need to query. 6 min ## [KERNEL: Self-Evolving Claude Code Configuration](https://medium.com/@ariaxhan/kernel-the-ultimate-self-evolving-claude-code-and-cursor-configuration-system-a3ddeb7f4d32) How I stopped fighting my config and let it learn instead. 6 min ## [This AI Analyzes My Entire Life](https://medium.com/@ariaxhan/the-synthesis-pool-0ce814fdfa5f) The Synthesis Pool: a personal AI that costs $0/month to run. 6 min ## [Claude vs. GPT vs. Gemini: You’re Benchmarking the Wrong Thing](https://medium.com/@ariaxhan/claude-vs-gpt-vs-gemini-youre-benchmarking-the-wrong-thing-762c7e9cca5d) These benchmarks are not only comparing models. They are comparing the systems built around those models and the design of the benchmarks themselves. 4 min ## [Opus 4.8 vs 4.7 vs Sonnet vs Haiku: When the Expensive Model Is Worth It](https://medium.com/@ariaxhan/opus-4-8-vs-4-7-vs-sonnet-vs-haiku-when-the-expensive-model-is-worth-it-44892a75d5c5) A new model dropped with impressive numbers. The only question that matters: will you feel any difference in the work you actually do? 12 min ## [What an AI Detector Actually Measures](https://medium.com/@ariaxhan/what-an-ai-detector-actually-measures-86b452979a5a) AI detectors promise to tell you if a machine wrote something. What they actually measure is much narrower, and shakier. 6 min ## [I Stopped Debugging Claude’s Code and Started Debugging My Prompts Instead](https://medium.com/@ariaxhan/how-i-stopped-debugging-ai-generated-code-1be33241b3c8) Debugging AI-generated code usually isn’t about fixing bad code. It’s about giving the model context it never had. 7 min ## [How to Make Claude Code Actually Work](https://medium.com/@ariaxhan/how-to-make-claude-code-actually-work-structure-memory-and-multi-agent-workflows-6d32b1d815d2) The most capable AI coding tool available. Also completely chaotic. 12 min ## [Stop Copying Other People's AI Setups. Build One That's Actually Yours.](https://medium.com/@ariaxhan/stop-copying-other-peoples-ai-setups-build-one-that-s-actually-yours-e1a05ebabc2a) Borrowed AI workflows aren't accountable to your work. Build one that's tested against your own evidence. 10 min ## [Automations with Claude Code](https://medium.com/@ariaxhan/automations-with-claude-code-personalized-proactive-emails-and-code-poetry-from-local-context-3a7e93bf5a3d) A pattern for proactive AI on your own machine. 4 min ## [From Friction to Flow: Building a Command Library](https://medium.com/@ariaxhan/from-friction-to-flow-building-a-command-library-for-claude-code-a9eb19f7dce2) Commands as cognitive offloading. Stop remembering, start invoking. 5 min ## [10 Things I Wish I Knew About AI Coding](https://medium.com/@ariaxhan/10-things-i-wish-i-knew-when-i-started-using-ai-for-coding-887c26a6c1d1) Hard-won lessons from daily production use of AI coding tools. 5 min ## [Engineering the Soul](https://medium.com/@ariaxhan/engineering-the-soul-49428c073c4e) We ask engineers to explain the ghost in the machine. The novelists have been documenting it for years. 6 min ## [I Tested OpenAI's New Codex Desktop App](https://medium.com/@ariaxhan/i-tested-openais-new-codex-desktop-app-the-ui-is-the-real-product-c2c59bdcb5f6) OpenAI shipped a genuinely novel interface. Then the model opened its mouth. 5 min ## [I Run 25 Websites, 10 Databases, and a Fleet of Apps for $0](https://medium.com/@ariaxhan/i-run-25-websites-10-databases-and-a-fleet-of-apps-for-0-27ec36756668) An honest inventory of everything I have running on the internet, and the free tiers that carry all of it. 9 min ## [What a Year of AI Taught Me About Freedom](https://medium.com/@ariaxhan/what-a-year-of-ai-taught-me-about-freedom-86b2bd4e31c8) There is a version of the future where AI makes corporations stronger, and one where it makes them irrelevant. 4 min # Let's talk I like technically ambitious work with thoughtful people, especially when the AI matters but the people matter more. ## Practical AI workflows I teach and build AI workflows for founders and independent builders: research, writing, operations, and decision-making systems they can keep using themselves. ## Internal operations workflows For companies, I build internal workflows that connect AI to the tools, files, and knowledge already in use. The work starts with the operation as it runs now, including where context gets lost and effort gets repeated. ## AI products for founders I work with founders and independent builders from an early idea through a working AI product. The implementation stays close to the person making the decisions, so the product can change as the problem becomes clearer. ## Agentic system architecture I built multi-agent coordination when it was still in its infancy: a single prompt fanning out into a family of agents in a coordinated dependency graph, each with its own role and tools, communicating through handoff notes instead of expensive cross-talk. I design structure so agent work becomes artifacts instead of fog. ## Evals, monitoring, and quality layers Checks that tell you whether the AI is doing the thing before a customer, teammate, or future version of you finds out the hard way. ## Claude Code / AI coding workflow hardening I spend a vast majority of my time talking to Claude Code. I can make the workflow calmer, more accountable, and less like a very expensive chaos machine, including the part nobody warns you about: agents behave very differently against years of legacy code than a project built from scratch. ## Memory, context, and knowledge systems Context is, in fact, everything, and it should not depend on whoever happens to remember it that week. I build the memory and knowledge layers that let people and agents hold their context: structured artifacts, richer recall, and transparency into exactly what the AI is referencing. ## AI product review and repair If an AI-assisted build mostly works but has become hard to debug, extend, or trust, I can review the product, trace where it is failing, and help turn the prototype into something you can keep building. Email: ariaxhan@gmail.com Book: https://cal.com/aria-han/15min