Your code doesn’t
leave the building.
Heap Code is a model-agnostic AI coding assistant for VS Code, your terminal, and your browser — chat, completions, inline edit, an autonomous agent, and AI PR review. Point it at Ollama on your own machine, or bring your own key to OpenAI, Groq, or any OpenAI-compatible endpoint. No account. No proprietary backend.
Extension v0.8.0 · CLI v0.7.0 · free for personal, internal & noncommercial use
Local: prompts, completions, and embeddings run entirely inside the boundary above — nothing crosses the line.
The terminal CLI
Node.js 20+ · macOS, Linux, Windows · no account, no sign-in
- Same engine as the extension. The agent, permissions, checkpoints, semantic search, and MCP support are shared code — not a reduced second implementation.
- /pr-review reviews your branch's PR file by file, then posts a real GitHub review with inline comments and one-click suggestions — only after you approve it. /pr-review deep adds a verification pass that drops false positives first.
- Headless mode. -p "task" runs the whole agent loop with no TTY — plain text or --json events, four permission modes for unattended runs, non-zero exit on failure. Built for CI.
- Every action gated. Same permission prompts as the editor, plus shadow-git checkpoints you can rewind step by step.
- Your models. Local Ollama or LM Studio, a box on your LAN, or any OpenAI-compatible endpoint — configured once, per role.
The browser UI
Runs on your machine · localhost by default · no account, no hosted service
- One command. heapcode web in any folder opens a browser tab pointed at that workspace — same agent, same permissions, same checkpoints as the CLI and the extension.
- A real workspace panel. Per-file diffs, one-click revert, a checkpoint timeline to rewind to, the file tree, streamed command output, and what the agent has actually indexed about your code.
- Things a terminal can’t do. Paste a screenshot straight into the message box. Get a desktop notification when a long run finishes in a background tab. Read rendered diagrams and HTML the agent built for you.
- Close the tab, keep the run. The server owns the work, not the browser — reopening reattaches to whatever is still going.
- Localhost by design. A fresh token per launch, never left in the address bar. --host opts in to reaching it from your phone, and says so loudly. Your API keys stay on the server and never reach the page.
New — Heap Chat
Ships in the CLI · heapcode chat [folder] · or at /chat under heapcode web
- Your files, not your code. Point it at a folder of PDFs, Word documents, spreadsheets and photos. It reads them, searches them, and answers from them — a scanned receipt is findable by what it says.
- Every answer shows its source. It says which file an answer came from, and a figure can be traced back to the line it was read from.
- It never writes over your files. A summary, a table or a report lands beside the conversation as a draft. Saving it into your folder is your call.
- Same engine, its own product. Its own tools, prompt, history and memory — sharing only the engine, the models you configured, and the index with Heap Code.
- Connectors with sign-in. MCP servers work here too, including hosted ones that need you to log in — and it asks before running one.
Also on this engine — heapbrowse
- A Chrome side panel that works the page you are on. Ask about it, or have it search, filter, compare across pages and fill in forms — and hand back the steps only you should take.
- Nothing happens without you seeing it. Every click or form fill is checked against what the element really is, and asks first. Banking and password sites are refused outright.
- Your own models. Point it at Ollama on your machine or your LAN — no server setup needed — and the pages you read never leave home.
Chrome extension · side panel · sites granted one at a time. Not on the Chrome Web Store yet: download the heapbrowse .zip from the latest release, unzip it, and choose “Load unpacked” in chrome://extensions.
Spec sheet
| Telemetry | Anonymous usage events only (feature usage, error counts) — never code, prompts, or file contents/paths. Opt-out via a setting. |
| Account required | No. No sign-in, no license server. |
| Model lock-in | None — any OpenAI-compatible endpoint, switch providers per-role (chat / edit / apply / completion / agent / embeddings / rerank). |
| Local model support | Yes — Ollama, LM Studio, vLLM, LocalAI. |
| Surfaces | VS Code extension, a standalone terminal CLI (heapcode), and a local browser UI (heapcode web) — one shared engine, so behaviour matches. Chat history, connections and project MCP servers are shared between them. Same engine: Heap Chat (heapcode chat) and the heapbrowse Chrome extension. |
| API key storage | Extension: OS keychain via VS Code SecretStorage. CLI: a chmod 600 file in ~/.heapcode — no keychain dependency, so headless and CI machines work. Never in settings files either way. |
| License / price | Free for personal, internal, research, nonprofit, education, and government use (PolyForm Noncommercial 1.0). |
Parts list — works with
Localruns on your machine
- Ollama
- LM Studio
- vLLM
- LocalAI
Cloudbring your own key
- OpenAI
- Azure OpenAI
- OpenRouter
- Groq
- Together AI
- NVIDIA NIM
Features
Chat
Streaming markdown, slash commands, @file/@selection/@workspace context, per-workspace history.
Completions
Ghost-text FIM tuned per model family, debounced and cancellable, repo-aware via the semantic index.
Inline edit
Select code, describe the change, review a native diff, accept from the title bar.
Agent mode
Reads, searches, edits, and runs commands autonomously. Shift+Tab cycles how much runs unattended — Plan (read-only, mutating tools aren’t even offered), Confirm, Auto-edit, or Auto — changeable mid-run, and Auto still stops for anything destructive. One-click revert on everything.
Semantic search
AST-aware chunking, hybrid embeddings + keyword search, degrades gracefully with no embedder configured.
MCP & tool interop
Register Model Context Protocol servers, including hosted ones that need a sign-in; each one gets only the environment variables it needs, not your whole shell. Other extensions' language-model tools show up in agent mode too.
Safety guardrails
Untrusted content — files, URLs, search results, MCP output — is tagged and quarantined before it reaches the model; hallucinated package installs are checked against the registry and blocked automatically; and agent web fetches can't reach private, loopback, or cloud-metadata addresses, at any redirect hop.
Granular checkpoints
Every tool call gets its own shadow-git snapshot — rewind any single step, not just the whole turn or the whole session.
Personas & Plan/Act
Architect, Debug, and Reviewer personas restrict which tools the agent is even offered; an optional plan-then-approve gate holds every change until you say go. Personas and permission modes compose — whichever is stricter wins.
Sub-agent delegation
Opt-in delegate_task hands a self-contained piece of work to an isolated sub-agent with its own context — never more permissive than the persona that spawned it.
AI PR review
Reviews a pull request file by file — never a truncated diff — then posts line-anchored comments with one-click suggestions via gh, after an explicit confirm. Deep mode adds a verification pass that drops false positives before you see them.
Web search
Off until you configure it. DuckDuckGo needs no key; Brave, Tavily and Serper take one, and a self-hosted SearXNG keeps queries off third-party infrastructure entirely. Results are quarantined as untrusted, same as any fetched page.
Runs on models that can’t tool-call
Plenty of local builds have no tool support in their chat template and reject the request outright. Heap Code notices, switches to a text-based tool protocol, and keeps going — so a model that would otherwise fail mid-task still finishes one.
Terminal CLI
npm i -g @heaplabs/heapcode-cli — the same agent, permissions, and checkpoints in your terminal, plus a headless -p mode with JSON output for CI.
Team bundles
Export or import a project's memory, instructions, and skills as a single file, so a team shares one setup.
Local audit dashboard
See exactly what ran, what was approved or denied, and what got reverted — computed entirely on your machine, nothing sent anywhere to produce it.