Skip to content

AI Agent Development 2026

AAD

Deep Dive · AI Agents

AI Agent Development 2026 — from the modern agent harness to a working build.

Artifact: a working agent (genie-claw) · Measure: reliability, latency, cost, safety

Parent: Phase 3 — Artificial Intelligence · Track B

Build applications on top of large language models — agents, RAG, tool use, GenAI products.

Prerequisites: Module 1 (Neural Networks), Module 2 (Frameworks — understand transformers and PyTorch).

Role targets: Agentic AI Engineer · GenAI Engineer · AI Engineer


A theory-first path: start from what a modern agent is in 2026, learn the fundamentals, build the core parts one by one, study a real harness (OpenClaw), then build your own (genie-claw). Lecture files keep their original numbers — this is the recommended reading order, not a renumbering. → Flat lecture index

Currency note: model names, context windows, SDK features, and pricing change fast. This course teaches the stable layer — model APIs, tool protocols, run loops, workflow graphs, gateways, telemetry, and policy boundaries. Always check provider docs before copying model IDs or prices into production.

Module 1 · Start here — what a modern agent is, and how to build one

The conceptual core of the whole course: what changed in 2026, the harness (the runtime around the model), and the two-part foundations of building an agent — model/tools/instructions, then orchestration and guardrails.

# Lecture
01 The modern AI agent in 2026 — what changed (start here)
02 What is an AI agent harness? — the runtime around the model
03 Building agents I — foundations (model, tools, instructions)
04 Building agents II — orchestration & guardrails

Module 2 · Fundamentals — the model underneath

How the model actually works, and how you talk to it.

# Lecture
05 LLM fundamentals for agents
06 LLM from scratch — model mechanics
07 Prompt engineering & structured output

Module 3 · Core building blocks (one by one)

Each part of an agent in dependency order: tools → memory → retrieval → orchestration/frameworks → multimodal → skills → evaluation.

# Lecture
08 Tool use & function calling
09 Structured tools beat computer use
10 Memory systems
11 RAG — ingestion & embeddings
12 RAG — retrieval & reranking
13 Vector stores & embedding model selection
14 Efficient local RAG stack
15 Agent architecture patterns (ReAct, plan-execute, reflexion)
16 LangGraph — stateful workflows
17 Agent SDKs & runtime APIs
18 OpenAI Agents SDK
19 Multi-agent systems
20 Multimodal sub-agents
21 Agent skills — workflow discipline
22 Agent skills eval
23 Evaluation & observability

Module 4 · Production discipline & runtime

Making an agent reliable, safe, and shippable: security, durable state, deterministic startup, runtime choice, lifecycle, and deployment.

# Lecture
24 Runtime discipline & AI runtime security
25 AI agent security engineer
26 Session as source of truth — event-sourced agent state
27 Deterministic startup
28 Runtime strategy — Node, Bun, Rust, and edge packaging
29 Agentic SDLC
30 Production deployment

Module 5 · Example — anatomy of a real harness (OpenClaw)

A production-style, local-first assistant taken apart piece by piece — the concepts of Modules 1–4 made concrete in one system.

# Lecture
31 Gateway architecture
32 Routing & sessions
33 Multi-agent isolation
34 Operations & security
35 The agent loop
36 Cron & scheduled agent runs
37 System prompt architecture
38 App SDK & typed RPCs
39 Gateway RPC protocol
40 OpenClaw threat model (MITRE ATLAS for agent security)
41 Pi — the minimal agent beneath OpenClaw

Module 6 · Practice — build genie-claw

The capstone: apply Modules 1–5 by building genie-claw — your own minimal agent harness that pairs a local LLM runtime (the GeniePod genie-ai-runtime) with an OpenClaw-style gateway: sessions, tools, guardrails, and a run loop you control end-to-end. The labs below are the stepping stones.

Lab Build
Lab 01 Research agent with tool use
Lab 02 Multi-agent code-review pipeline
Lab 03 Production RAG system
Lab 04 TokenJuice output compaction
Lab 05 OpenMeow App SDK dogfood (macOS)
Lab 06 · Capstone genie-claw — build your own minimal agent harness end-to-end

The inference- and kernel-leaning deep dives (GPU kernels, FP8 KV-cache, performance tracing, sequence parallelism, small-MoE reasoning) used to live here as an appendix. They belong with ML Systems Engineering, so they now live in the MLSys Deep Dives collection (Phase 5), alongside the AI Inference Engineer 2026 course.


Why This Matters for AI Hardware

Agentic AI creates the inference demand that drives chip design: - Long-context attention (128K+ tokens) → L5: HBM bandwidth, memory hierarchy - Multi-turn tool calling → L3: low-latency kernel launch, stream scheduling - Batch inference serving → L2: TensorRT-LLM, in-flight batching optimization - RAG vector search → L1: cuVS/FAISS acceleration on GPU

Understanding these workloads helps L2 (compiler) and L5 (architecture) engineers design for real usage patterns.


The strongest application trend is that AI is moving from chat interfaces to persistent, tool-using systems that act inside real workflows.

Two useful reference patterns are: - agentic coding systems such as Claude Code - local-first personal assistant systems such as OpenClaw

These are worth studying because they show what modern AI applications actually look like in production-like usage, not just in demos.

1. Coding Agents Are Replacing Single-Step “Code Completion”

This is the clearest shift.

Tools like Claude Code are no longer limited to autocomplete in an editor. They act more like software workers: - read and navigate a repository - plan changes across multiple files - run tests and verification loops - create commits and pull requests - connect to external tools through MCP - load reusable skills, hooks, and plugins

Anthropic’s official docs describe this directly: Claude Code can automate routine engineering work, work with git, connect tools through MCP, spawn multiple agents, and run in CI/CD workflows. Its public repository and plugin system make it a good reference for how coding agents are becoming a real application category rather than a novelty.

Why this matters for hardware: - coding agents create long-running, tool-rich inference sessions instead of short chat turns - they increase demand for low-latency iteration loops, larger context windows, and higher background inference volume - they push AI products into developer infrastructure, where reliability, permissioning, and automation matter as much as model quality

2. Personal AI Is Moving Toward Local-First Control Planes

OpenClaw is a useful reference for a different trend: AI assistants that are not “one web page with one chat box,” but a control plane that stays running and connects many surfaces.

From the official OpenClaw repo and docs: - one long-lived local Gateway owns channels, sessions, tools, and events - the assistant can operate across WhatsApp, Telegram, Slack, Discord, WebChat, and other channels - it supports voice, mobile nodes, live canvas UI, and multi-agent routing - it treats inbound messages as untrusted input and documents a concrete security model

This shows where application design is going: - persistent assistants instead of one-off prompts - multi-channel delivery instead of one frontend - local or operator-controlled infrastructure instead of only cloud-hosted chat - agents as routed services with isolated workspaces and memory

Why this matters for hardware: - always-on assistants create steady inference demand, not just bursty usage - multimodal assistants increase pressure on device memory, streaming, and local inference paths - local-first designs make edge hardware, Jetson-class devices, mobile nodes, and hybrid cloud/edge deployment more relevant

3. MCP, Plugins, and Hooks Are Becoming the Real Application Surface

Another major trend is that the model alone is no longer the full product.

Modern AI systems are increasingly defined by: - MCP connectors to external systems - plugins for reusable workflows - hooks and automations around model actions - skills and custom agents for domain-specific behavior

Claude Code’s docs explicitly position MCP, plugins, skills, hooks, monitors, and custom agents as first-class extension surfaces. OpenClaw similarly treats tools, plugins, channels, nodes, and gateway protocols as part of the product architecture.

The implication is important: the application layer is becoming a tool-and-protocol ecosystem, not just a prompt template.

4. Agent SDKs Are Becoming Runtime Layers, Not Just API Wrappers

Modern agent SDKs now sit above raw model calls. They increasingly manage: - agent loops - tool dispatch - handoffs between specialists - sessions and state - guardrails and human review - tracing and evaluation - MCP server integration

This matters because production agent systems need more than a messages.create() call. They need a repeatable runtime contract: what tools are available, which identity executes them, which actions require approval, where state is stored, and what gets logged.

The practical design rule for this course is:

Provider API details belong in adapters.
Product behavior belongs in your runtime contract.
Security decisions belong outside the LLM.

This is why Lecture 17 teaches SDKs and runtime APIs as a general layer instead of treating one vendor SDK as the architecture.

5. Multi-Agent Structure Is Becoming Practical, Not Theoretical

The industry has moved beyond “one model, one prompt, one answer.”

Current systems increasingly use: - a lead agent plus worker agents - isolated workspaces per task or user - explicit routing rules - background monitors and event-driven triggers

Claude Code exposes multiple-agent workflows and custom agents. OpenClaw exposes multi-agent routing with isolated workspaces, session stores, and bindings from inbound channels to specific agents.

This is a practical trend because it maps well to real products: - support workflows - developer workflows - personal assistant workflows - project automation and governance workflows

6. Security and Permission Boundaries Are Now Core Product Features

This is one of the biggest changes from early LLM apps.

Current AI applications increasingly ship with: - pairing and allowlists - tool permission boundaries - gateway auth and signed connections - sandboxing - safe output policies - moderation and prompt-injection defenses

OpenClaw’s docs emphasize pairing approval, DM safety, sandbox modes, and gateway auth. Claude Code’s plugin system and tool model emphasize explicit structure, scoped extensions, and operational safety. GitHub’s agentic workflow material also frames security, permissions, and isolated execution as central rather than optional.

This means “application architecture” now includes trust boundaries, not just prompts and UI.

7. What To Learn From These Examples

Do not study OpenClaw and Claude Code just because they are popular. Study them because they represent two high-signal application patterns:

  • Claude Code pattern: AI embedded into developer workflows, repositories, CI, tools, and review loops
  • OpenClaw pattern: AI embedded into messaging, voice, mobile nodes, control planes, and personal automation

Together they show that the current application frontier is: - agentic - tool-connected - persistent - permissioned - multi-surface - operationally observable

For this roadmap, the important takeaway is that Track B should teach the workloads and architectures that real AI products are converging toward, because those products are what downstream systems and hardware will ultimately serve.


1. LLM Fundamentals for Engineers

  • Transformer architecture: attention mechanism, KV-cache, positional encoding
  • Tokenization: BPE, SentencePiece, vocabulary size impact on embedding layer
  • Inference mechanics: prefill (compute-bound) vs decode (memory-bound), autoregressive generation
  • Scaling laws: parameter count vs dataset size vs compute budget

2. Agentic AI

  • What agents are: LLM + tools + memory + planning loop
  • Agent runtimes and frameworks: raw provider APIs, OpenAI Agents SDK, LangGraph, MCP servers, OpenClaw-style gateways, CrewAI/AutoGen for experiments
  • Tool use: function calling, API integration, code execution
  • Security boundaries: prompt injection, tool abuse, least-privilege credentials, human approval for risky actions
  • Memory: conversation history, vector store retrieval, working memory
  • Planning: chain-of-thought, ReAct, tree-of-thought, self-reflection
  • Multi-agent systems: task decomposition, agent collaboration, orchestration

Projects: 1. Build an agent that uses tools (web search, calculator, code execution) to answer complex questions. 2. Build a multi-step research agent: given a topic, search → synthesize → write a report.


3. RAG (Retrieval-Augmented Generation)

  • Architecture: document ingestion → chunking → embedding → vector store → retrieval → generation
  • Embedding models: sentence-transformers, OpenAI embeddings, Cohere
  • Vector stores: FAISS, Chroma, Pinecone, Weaviate, Milvus
  • Chunking strategies: fixed-size, recursive, semantic, document-structure-aware
  • Retrieval: similarity search, hybrid (dense + sparse), re-ranking
  • RAG security: treat documents as untrusted input, screen uploads and retrieved chunks, defend against indirect prompt injection
  • Evaluation: faithfulness, relevance, hallucination detection

Projects: 1. Build a RAG pipeline over a technical documentation corpus. Evaluate retrieval quality. 2. Compare FAISS (CPU) vs FAISS (GPU) vs cuVS for vector search latency at 1M documents.


4. GenAI Product Development

  • Prompt engineering: system prompts, few-shot, chain-of-thought, output formatting
  • Fine-tuning: LoRA, QLoRA, full fine-tuning on domain data
  • Evaluation: automated metrics (ROUGE, BLEU), LLM-as-judge, human evaluation
  • Guardrails: input moderation, output filtering, content safety, prompt-injection defense, hallucination mitigation
  • Production deployment: API design, streaming, rate limiting, cost management
  • Runtime discipline: live telemetry, tool-call policy gates, least-privilege agent identities, audit trails, and runtime incident response
  • Deterministic startup: startup contracts, readiness checks, prompt/tool/policy versioning, memory hydration, and reproducible agent boot
  • Agent control planes: gateways, sessions, routing, and multi-surface delivery
  • Scheduled agent execution: cron jobs, isolated run sessions, delivery fallback, failure routing, retries, run logs, and retention
  • OpenClaw-style agent development: multi-agent isolation, workspace design, pairing, and operations for persistent local-first assistants

Projects: 1. Fine-tune a 7B model with QLoRA on a domain-specific dataset. Measure improvement vs base model. 2. Deploy a GenAI application with streaming, safety guardrails, and cost tracking. 3. Build a secure agent or RAG pipeline with prompt-attack detection, tool constraints, and output validation.


Resources

Resource What it covers
LangChain Documentation Agent and RAG framework
LangGraph Documentation Durable, stateful agent workflows, human-in-the-loop, memory, and tracing
OpenAI Agents SDK Agent loops, tools, handoffs, guardrails, sessions, tracing, and MCP integration
OpenAI API Agents Guide Current OpenAI guidance for code-first agent apps, tools, orchestration, and observability
Model Context Protocol Specification Standard protocol for tools, resources, prompts, hosts, clients, servers, and safety considerations
Claude Code Overview Agentic coding workflows, MCP, multi-agent use, and CI patterns
Claude Code Plugins Skills, agents, hooks, MCP servers, plugin structure, and distribution
Claude Code Repository Public implementation surface, examples, plugins, and project layout
Anthropic Cookbook Practical Claude API examples
OpenClaw Repository Local-first assistant architecture, channels, gateway model, and security defaults
OpenClaw Gateway Architecture Long-lived gateway, WS protocol, nodes, pairing, and remote access model
OpenClaw Features Multi-agent routing, media, channels, tools, apps, and provider support
GitHub Agentic Workflows Official GitHub framing for agentic CI/CD, permissions, and safe outputs
OWASP Top 10 for LLM Applications Prompt injection, insecure output handling, plugin/tool risk, excessive agency, and LLM app security
NIST AI RMF Generative AI Profile Governance and risk-management framing for generative AI systems
RAG best practices LlamaIndex documentation
Build a Large Language Model (From Scratch) (Raschka) LLM internals

Next

→ Module 4B — ML Engineering & MLOps