Ultimate Unity + VS Code Environment

Mastering your workspace visibility, Git tracking, and AI Indexing rules

- - - -

If you've ever opened a raw Unity project inside Visual Studio Code, you know exactly what happens: absolute chaos. Your file explorer gets immediately clogged with thousands of .meta files, your Git client tries to commit massive 3D models, and if you're using modern AI tools like Zoo Code, the indexer completely chokes trying to read baked lighting data.

Unity is a heavy engine that generates massive amounts of proprietary files. To keep your workflow clean, fast, and professional, you need to set boundaries on three different layers:

  • Git (.gitignore): Tells version control what shouldn't go to the cloud.
  • VS Code Explorer (.code-workspace): Tells your IDE what to hide from your human eyes.
  • AI Indexing (.rooignore): Tells AI agents like Zoo Code what files to ignore so you don't burn through tokens and memory.

We've done the deep research into the official Unity 6 documentation to compile the absolute most exhaustive lists of 3D, audio, video, and asset extensions you will ever need. Let's get your project set up properly.

1. The Git Boundary (.gitignore)

Your repository should only contain source code, raw text-based configuration, and essential project settings. It should never track local builds, temporary cache, or heavy local IDE settings. Place this file exactly at the root of your Unity project.

The Golden Rule of Unity Git

Never, ever ignore .meta files inside your Assets/ folder. Unity uses these as GUIDs to link scripts to GameObjects. If you don't commit a meta file, your teammates will open the project and find missing script references everywhere. Our script handles this correctly at the top.

๐Ÿ“„ .gitignore
# Based on github/gitignore Unity template + Unity 6 full asset coverage
# CRITICAL: .meta files inside Assets/ must NEVER be ignored
!/[Aa]ssets/**/*.meta

# Unity generated folders
/[Ll]ibrary/
/[Tt]emp/
/[Oo]bj/
/[Bb]uild/
/[Bb]uilds/
/[Ll]ogs/
/[Uu]ser[Ss]ettings/
/[Mm]emoryCaptures/
/[Rr]ecordings/
ExportedObj/

# IDE & Cache
.vs/
.idea/
.gradle/
.consulo/
*.csproj
*.unityproj
*.sln
*.suo
*.tmp
*.user
*.userprefs
*.pidb
*.booproj
*.svd
*.pdb
*.mdb
*.opendb
*.VC.db
*.pidb.meta
*.pdb.meta
*.mdb.meta
/[Aa]ssets/Plugins/Editor/JetBrains*

# Crash reports
sysinfo.txt
crashlytics-build.properties

# โ”€โ”€ 3D Model formats โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€
*.fbx
*.obj
*.dae
*.3ds
*.max
*.blend
*.c4d
*.ma
*.mb
*.skp
*.jas
*.lwo
*.lws
*.ltm
*.stl
*.glb
*.gltf
*.abc
*.spm
*.sbsar
*.st9
*.ztl
*.zpr
*.mud

# โ”€โ”€ Texture / Image formats โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€
*.psd
*.psb
*.png
*.jpg
*.jpeg
*.tga
*.tif
*.tiff
*.bmp
*.gif
*.hdr
*.exr
*.dds
*.ktx
*.ktx2
*.astc
*.pvr
*.svg
*.avif
*.apng
*.cur
*.iff
*.pict
*.raw
*.webp

# โ”€โ”€ Audio formats โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€
*.wav
*.mp3
*.ogg
*.aiff
*.aif
*.flac
*.xm
*.mod
*.it
*.s3m
*.aac
*.m4a
*.wma
*.opus

# โ”€โ”€ Video formats โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€
*.mp4
*.mov
*.avi
*.mkv
*.webm
*.flv
*.m4v
*.h264
*.h265
*.wmv
*.asf
*.ogv
*.mpeg
*.mpg

# โ”€โ”€ Font formats โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€
*.ttf
*.otf
*.fon

# โ”€โ”€ SpeedTree โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€
*.spm
*.st
*.st9

# โ”€โ”€ Addressables / Streaming โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€
/[Aa]ssets/[Aa]ddressable[Aa]ssets[Dd]ata/*/*.bin*
/[Aa]ssets/[Ss]treamingAssets/aa.meta
/[Aa]ssets/[Ss]treamingAssets/aa/*

# โ”€โ”€ Build outputs โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€
*.apk
*.aab
*.unitypackage
*.app
*.exe
*.x86
*.x86_64
*.bc

# โ”€โ”€ GitNexus local knowledge graph โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€
# Auto-generated by `gitnexus analyze`  -  machine-specific, do not commit
.gitnexus/

2. The Visual Boundary (VS Code Explorer)

You need to commit .meta files, but you definitely don't want to look at them while you're writing C# code. By using a .code-workspace file, we can utilize VS Code's files.exclude setting. This leaves the files perfectly intact on your hard drive, but renders them completely invisible in the VS Code sidebar.

Pro Tip

If you're using a Multi-Root Workspace (working across multiple Git repos in one window), this file makes managing them seamless. Just save this JSON as MyGame.code-workspace in your root folder and double-click it to open your project.

๐Ÿ“„ MyGame.code-workspace
{
  "folders": [
    { "path": "." }
  ],
  "settings": {
    "git.autoRepositoryDetection": "subFolders",
    "files.exclude": {
      "**/*.meta": true,

      "Library/": true,
      "Temp/": true,
      "Obj/": true,
      "Build/": true,
      "Builds/": true,
      "Logs/": true,
      "UserSettings/": true,
      "MemoryCaptures/": true,
      "Packages/": true,

      "**/*.csproj": true,
      "**/*.sln": true,
      "**/*.suo": true,
      "**/*.tmp": true,
      "**/*.user": true,
      ".vs/": true,

      "**/*.fbx": true, "**/*.obj": true, "**/*.dae": true,
      "**/*.3ds": true, "**/*.blend":true, "**/*.ma": true,
      "**/*.mb": true, "**/*.max": true, "**/*.c4d": true,
      "**/*.stl": true, "**/*.glb": true, "**/*.gltf": true,
      "**/*.abc": true, "**/*.skp": true, "**/*.sbsar":true,
      "**/*.st9": true, "**/*.ztl": true,

      "**/*.psd": true, "**/*.psb": true, "**/*.png": true,
      "**/*.jpg": true, "**/*.jpeg": true, "**/*.tga": true,
      "**/*.tif": true, "**/*.tiff": true, "**/*.bmp": true,
      "**/*.gif": true, "**/*.hdr": true, "**/*.exr": true,
      "**/*.dds": true, "**/*.ktx": true, "**/*.ktx2": true,
      "**/*.svg": true, "**/*.webp": true, "**/*.avif": true,
      "**/*.astc": true, "**/*.pvr": true,

      "**/*.wav": true, "**/*.mp3": true, "**/*.ogg": true,
      "**/*.aiff": true, "**/*.aif": true, "**/*.flac": true,
      "**/*.aac": true, "**/*.m4a": true, "**/*.opus": true,
      "**/*.xm": true, "**/*.mod": true,

      "**/*.mp4": true, "**/*.mov": true, "**/*.avi": true,
      "**/*.mkv": true, "**/*.webm": true, "**/*.flv": true,
      "**/*.m4v": true, "**/*.wmv": true, "**/*.ogv": true,

      "**/*.ttf": true, "**/*.otf": true,

      "**/*.asset": true,
      "**/*.unity": true,
      "**/*.prefab": true,
      "**/*.mat": true,
      "**/*.anim": true,
      "**/*.controller": true,
      "**/*.overrideController": true,
      "**/*.playable": true,
      "**/*.signal": true,
      "**/*.physicMaterial": true,
      "**/*.physicsMaterial2D": true,
      "**/*.guiskin": true,
      "**/*.fontsettings": true,
      "**/*.flare": true,
      "**/*.lighting": true,
      "**/*.cubemap": true,
      "**/*.rendertexture": true,
      "**/*.shadervariants": true,
      "**/*.spriteatlas": true,
      "**/*.terrainlayer": true,
      "**/*.mask": true
    }
  }
}

3. The AI Boundary (.rooignore)

If you use AI coding assistants like Zoo Code with Qdrant vector database indexing, you absolutely must restrict what it reads. If you don't, the AI will try to tokenize and index gigabytes of raw binary FBX files and MP3s. Not only will this cause a "fetch failed" timeout error, it actively poisons the AI's context window.

Place .rooignore in the same root folder as your .gitignore. Unlike Git, Zoo Code doesn't care about .meta files for tracking purposes, so we can happily block those from the AI entirely.

๐Ÿ“„ .rooignore
# โ”€โ”€ Unity generated โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€
Library/
Temp/
Obj/
Build/
Builds/
Logs/
UserSettings/
MemoryCaptures/
Recordings/
Packages/
Package Cache/
ExportedObj/

# โ”€โ”€ IDE โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€
.vs/
.idea/
*.csproj
*.sln
*.suo
*.user
*.userprefs

# โ”€โ”€ 3D Models โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€
*.fbx
*.obj
*.dae
*.3ds
*.blend
*.ma
*.mb
*.max
*.c4d
*.stl
*.glb
*.gltf
*.abc
*.skp
*.sbsar
*.st9
*.ztl
*.zpr

# โ”€โ”€ Textures & Images โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€
*.psd
*.psb
*.png
*.jpg
*.jpeg
*.tga
*.tif
*.tiff
*.bmp
*.gif
*.hdr
*.exr
*.dds
*.ktx
*.ktx2
*.svg
*.webp
*.avif
*.astc
*.pvr
*.apng
*.cur

# โ”€โ”€ Audio โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€
*.wav
*.mp3
*.ogg
*.aiff
*.aif
*.flac
*.aac
*.m4a
*.opus
*.xm
*.mod
*.it
*.s3m
*.wma

# โ”€โ”€ Video โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€
*.mp4
*.mov
*.avi
*.mkv
*.webm
*.flv
*.m4v
*.wmv
*.ogv
*.mpeg
*.mpg
*.h264
*.h265

# โ”€โ”€ Fonts โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€
*.ttf
*.otf
*.fon

# โ”€โ”€ Unity native asset types โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€
*.meta
*.asset
*.unity
*.prefab
*.mat
*.anim
*.controller
*.overrideController
*.playable
*.signal
*.physicMaterial
*.physicsMaterial2D
*.guiskin
*.fontsettings
*.flare
*.lighting
*.cubemap
*.rendertexture
*.shadervariants
*.spriteatlas
*.terrainlayer
*.mask
*.brush
*.inputactions
*.uss
*.uxml

# โ”€โ”€ Build outputs โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€
*.apk
*.aab
*.unitypackage
*.app
*.exe
*.dll
*.so
*.dylib

# โ”€โ”€ Misc binary / data โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€
*.bin
*.bytes
*.zip
*.tar
*.gz
*.7z

# โ”€โ”€ GitNexus graph database โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€
# Binary KuzuDB files  -  Zoo Code must use MCP tools to query these, never read directly
.gitnexus/

After implementing the .rooignore, remember to restart VS Code and click Clear Index Data in the Zoo Code settings before triggering a fresh codebase scan. Your indexing will be lightning fast, targeting strictly C# scripts and readable configs.

3b Zoo Code Setup

Zoo Code is the community-maintained open-source successor to Roo Code, developed at Zoo-Code-Org/Zoo-Code. It is a direct continuation of the same extension - same .roo/ folder structure, same .roomodes, same .rooignore, same MCP config format - with continued model updates, bug fixes, and community features. Install it from the VS Code Marketplace as ZooCodeOrganization.zoo-code.

Migrating from Roo Code

All your existing .roomcp.json, .roomodes, .rooignore, and .roo/skills/ files work with Zoo Code without any changes. Uninstall the old Roo Code extension, install Zoo Code, and reload VS Code. Your configuration carries over automatically. The official migration guide is at docs.zoocode.dev/roo-to-zoo-migration.

Add the Behavioral Rules File

Zoo Code loads .roo/rules/rules.md as a persistent system-level instruction file that applies to every mode in every session. This is the correct place for universal behavioral guidelines - the rules that should govern how the AI reasons and acts regardless of whether it is in Code, Architect, Debug, or a custom mode. Create the file now:

Create the folder first

Run this once in your project root terminal before creating the file:

Terminal - Windows PowerShell
mkdir -Force .roo\rules

Then create the file .roo\rules\rules.md with the content below. This is a set of first-principles behavioral constraints that reduce common LLM coding mistakes - favoring caution over speed, surgical edits over sweeping rewrites, and explicit clarification over silent assumptions. Credits: https://github.com/multica-ai/andrej-karpathy-skills/blob/main/CLAUDE.md

.roo\rules\rules.md
# CLAUDE.md

Behavioral guidelines to reduce common LLM coding mistakes. Merge with project-specific instructions as needed.

**Tradeoff:** These guidelines bias toward caution over speed. For trivial tasks, use judgment.

## 1. Think Before Coding

**Don't assume. Don't hide confusion. Surface tradeoffs.**

Before implementing:

- State your assumptions explicitly. If uncertain, ask.
- If multiple interpretations exist, present them - don't pick silently.
- If a simpler approach exists, say so. Push back when warranted.
- If something is unclear, stop. Name what's confusing. Ask.

## 2. Simplicity First

**Minimum code that solves the problem. Nothing speculative.**

- No features beyond what was asked.
- No abstractions for single-use code.
- No "flexibility" or "configurability" that wasn't requested.
- No error handling for impossible scenarios.
- If you write 200 lines and it could be 50, rewrite it.

Ask yourself: "Would a senior engineer say this is overcomplicated?" If yes, simplify.

## 3. Surgical Changes

**Touch only what you must. Clean up only your own mess.**

When editing existing code:

- Don't "improve" adjacent code, comments, or formatting.
- Don't refactor things that aren't broken.
- Match existing style, even if you'd do it differently.
- If you notice unrelated dead code, mention it - don't delete it.

When your changes create orphans:

- Remove imports/variables/functions that YOUR changes made unused.
- Don't remove pre-existing dead code unless asked.

The test: Every changed line should trace directly to the user's request.

## 4. Goal-Driven Execution

**Define success criteria. Loop until verified.**

Transform tasks into verifiable goals:

- "Add validation" โ†’ "Write tests for invalid inputs, then make them pass"
- "Fix the bug" โ†’ "Write a test that reproduces it, then make it pass"
- "Refactor X" โ†’ "Ensure tests pass before and after"

For multi-step tasks, state a brief plan:

```
1. [Step] โ†’ verify: [check]
2. [Step] โ†’ verify: [check]
3. [Step] โ†’ verify: [check]
```

Strong success criteria let you loop independently. Weak criteria ("make it work") require constant clarification.

---

**These guidelines are working if:** fewer unnecessary changes in diffs, fewer rewrites due to overcomplication, and clarifying questions come before implementation rather than after mistakes.
Rules File Scope - Global vs. Mode-Specific
  • .roo/rules/rules.md - applies to all modes (Code, Architect, Debug, Ask, and custom modes). Use for universal behavioral constraints like the ones above.
  • .roo/rules-code/rules.md - applies to Code mode only. Use for language-specific or project-specific coding standards that only make sense during implementation.
  • .roo/rules-architect/rules.md - applies to Architect mode only. Use for system design conventions, ADR formats, or diagram standards.
  • Multiple .md files can coexist in the same rules folder - Zoo Code loads them all.
Commit This File

.roo/rules/rules.md contains no secrets and no machine-specific paths. Commit it to Git so every team member automatically gets the same behavioral constraints when they clone the repo.

3c Ponytail for Zoo Code

Ponytail (DietrichGebert/ponytail) is an open-source agent skill that puts a lazy senior developer inside Zoo Code. Before writing any code, the agent stops at the first rung that holds: Does this need to exist? Does it already exist in this codebase? Can stdlib or a native platform feature cover it? Is an installed dependency enough? Can it be one line? Only then: the minimum that works. Published agentic benchmarks across 12 feature tasks (FastAPI + React repo, Haiku 4.5) show โˆ’54% LOC, โˆ’20% cost, โˆ’27% time with 100% safety maintained - validation, error handling, security, and accessibility are never cut (source: repo benchmarks). It is not "write fewer tokens" - it is "only write what the task needs." 49K+ GitHub stars, MIT licensed.

A Note on Caveman

If you previously explored Caveman (JuliusBrussee/caveman), note the difference in approach: Caveman focuses on compressing output verbosity - it reduces what the model prints (~65% output token savings, source). Ponytail takes a different path - it prevents unnecessary code from being written in the first place: โˆ’54% LOC, โˆ’22% tokens, โˆ’20% cost, โˆ’27% time (source). Both can coexist, but Ponytail's rules-file integration below is the simpler setup for Zoo Code.

What Ponytail Actually Does

Ponytail installs a simple, persistent rule file that forces the agent to climb a 7-rung decision ladder before writing any code. The ladder stops at the first rung that solves the task:

  1. YAGNI: Does this even need to exist? Skip it.
  2. Already in this codebase? Reuse it - don't rewrite it.
  3. Standard library? Use that.
  4. Native platform feature? Use that.
  5. Already an installed dependency? Use that.
  6. Can it be one line? Write one line.
  7. Only then: Write the minimum code that works.

Ponytail ships with these slash commands:

Skill / Tool What It Does
/ponytail [lite|full|ultra|off] Sets intensity for the session. full is the default. ultra is for codebases with significant over-engineering debt.
/ponytail-review Reviews the current diff for over-engineering; returns a delete-list of lines that can be simplified or removed.
/ponytail-audit Audits the entire repo for over-engineering - not just the current diff. Finds duplicate implementations, unnecessary abstractions, and unused code paths.
/ponytail-debt Harvests ponytail: shortcut comments scattered through the codebase into a single tracked ledger of simplification opportunities.
/ponytail-gain Shows the measured impact scoreboard (LOC, cost, speed) from the published benchmark for your current session.
Lazy, Not Negligent

Ponytail's rules explicitly preserve validation, error handling, security, and accessibility. "Lazy" means efficient - the best code is the code never written, but the code that is written must be correct on edge cases. Shortest working diff wins. Boring beats clever.

Installation & Setup

Ponytail integrates with Zoo Code through the same .roo/rules/ mechanism documented in Section 3b. Zoo Code loads all .md files in .roo/rules/ as persistent system instructions. The canonical Ponytail rule file lives at .clinerules/ponytail.md in the repo - Zoo Code (inheriting Roo Code's rule engine) reads this as .roo/rules/ponytail.md.

Prerequisite - .roo/rules/ must exist

If you followed Section 3b, this directory already exists with rules.md inside it. If not, run mkdir -Force .roo\rules (Windows) or mkdir -p .roo/rules (macOS/Linux) first.

๐Ÿ–ฅ๏ธ Windows PowerShell - Fetch Ponytail rule file
Invoke-WebRequest `
  -Uri "https://raw.githubusercontent.com/DietrichGebert/ponytail/main/.clinerules/ponytail.md" `
  -OutFile ".roo\rules\ponytail.md"
๐Ÿ–ฅ๏ธ macOS / Linux - Fetch Ponytail rule file
curl -o .roo/rules/ponytail.md \
  https://raw.githubusercontent.com/DietrichGebert/ponytail/main/.clinerules/ponytail.md

That's it. Zoo Code discovers ponytail.md automatically on the next session - no registration, no config files, no plugin marketplace. The canonical .clinerules/ponytail.md is ~500 words, already compact, and requires no separate compression step.

Mode Scoping (Optional)

By default, .roo/rules/ponytail.md applies to all modes - Code, Architect, Debug, Ask, and any custom modes. To restrict Ponytail to Code mode only, move the file to .roo/rules-code/ponytail.md. The recommendation is to keep it global: the ladder applies equally during debugging (fix root cause, not symptom) and Ask-mode responses (no example boilerplate nobody asked for). Architect mode benefits from YAGNI rung 1 - question whether the plan requires every proposed component.

Commit This File

.roo/rules/ponytail.md contains no secrets and no machine-specific paths. Commit it to Git so every teammate automatically gets the same anti-over-engineering constraints when they clone the repo - exactly like the existing rules.md.

Interaction with Zoo Code Modes

Ponytail's rule file coexists with the existing .roo/rules/rules.md (Karpathy behavioral rules from Section 3b). They are complementary: rules.md governs how the AI reasons (caution, surgical edits, clarify before coding), while Ponytail governs how much code it writes (climb the ladder before building anything). Both files are loaded simultaneously - no conflict.

Mode Role What Ponytail Does There
architect (built-in) Planning & system design YAGNI rung applies - question whether the plan requires every proposed component. Ponytail supplements, not replaces, architectural reasoning.
code (built-in or custom) Implementation & refactoring Full 7-rung ladder before every file edit. This is where Ponytail has its largest impact - preventing npm installs for things the browser already does natively, avoiding unnecessary abstractions, and writing the minimum that works.
debug (built-in) Root-cause analysis "Fix the root cause, not the symptom" directive - grep every caller of the function you touch and fix the shared function once (the smaller diff). On a shared-helper bug-fix trap, baseline fixes root cause 1/6; Ponytail does 6/6 (PR #253).
ask (built-in) Explanation & review Keeps answers minimal - no example boilerplate nobody asked for. Prefers one concise example over a gallery.

Common Mistakes

  • Placing ponytail.md in .rooignore. Zoo Code must read rule files to function. Never block .roo/rules/ from AI indexing - rule files are text, not binary assets. The existing .rooignore in Section 3 correctly excludes only binary/generated files.
  • Confusing Ponytail's ladder with "write fewer lines at any cost." The rule explicitly preserves validation, error handling, security, and accessibility. Lazy means efficient, not reckless.
  • Using Ponytail in Architect mode instead of running the ladder within Architect. Ponytail supplements architectural reasoning - it does not replace it. The YAGNI rung questions scope, not structure.
  • Forgetting to update ponytail.md after upstream releases. The Ponytail repo evolves. Periodically re-fetch the latest .clinerules/ponytail.md and diff it against your local copy. The repo includes a check-rule-copies.js script to verify alignment.
  • Assuming Ponytail needs a separate compression step. Unlike Caveman's caveman-compress workflow, the ponytail.md rule file is already compact (~500 words). No additional compression tooling is needed - lower context overhead out of the box.
  • Not committing ponytail.md to Git. Rule files contain no secrets. Commit them so every teammate gets the same anti-over-engineering constraints automatically on clone - identical to the guidance for rules.md in Section 3b.
Final Recommendation

For Zoo Code users: (1) fetch ponytail.md into .roo/rules/ - one file, one command, zero ongoing maintenance beyond periodic upstream sync; (2) it coexists cleanly with the existing rules.md behavioral file; (3) it applies across all Zoo Code modes without requiring custom mode creation; (4) it prevents over-engineering at the source. If you previously set up Caveman custom modes (caveman-code, caveman-debug, caveman-ask), they remain functional - the Ponytail rule file runs alongside them without conflict.

4. Graph-Aware AI (GitNexus + Zoo Code)

Standard AI coding assistants search your codebase using text similarity - they find code that looks related but cannot tell you what actually calls what. GitNexus solves this by building a local knowledge graph of your entire codebase using Tree-sitter AST parsing, then exposing that graph to Zoo Code via the Model Context Protocol (MCP). Before Zoo Code writes a single line, it already knows the blast radius of every change - every caller, every dependent execution flow, every file that could break.

This is a completely local, zero-server setup. Your source code never leaves your machine. Even when using an external LLM for reasoning, only the specific tool output (the graph query result) is transmitted - not the full repository.

What You Get

Once set up, Zoo Code gains 7 specialized MCP tools - blast radius analysis, graph-aware multi-file rename, execution flow tracing, and more - all running locally via a KuzuDB-backed knowledge graph. Zoo Code is not on GitNexus's official auto-setup list (that's Claude Code, Cursor, Windsurf, OpenCode), so this guide shows you the correct manual path.

Prerequisites

  • Node.js v18 or later installed globally
  • Zoo Code VS Code extension installed and active
  • A project with a Git repository initialized at its root
  • .gitnexus/ already added to your .gitignore and .rooignore (covered in Sections 1 & 3 above)

Step 1 - Install GitNexus & Index Your Project

Run the following from the root of your project. Install the CLI globally first to avoid version lag on repeated runs, then trigger the full indexing pipeline:

๐Ÿ–ฅ๏ธ Terminal
npm install -g gitnexus
gitnexus analyze

This triggers the full multi-phase indexing pipeline: file tree walking โ†’ symbol extraction via Tree-sitter AST (functions, classes, interfaces, variables) โ†’ cross-file reference resolution (imports, call sites, definitions) โ†’ module clustering via Leiden community detection โ†’ execution flow tracing from entry points to leaf functions. The output is a .gitnexus/ directory containing the local KuzuDB knowledge graph. It also auto-generates AGENTS.md and CLAUDE.md at your project root - human-readable codebase overview files for AI agents that are safe to commit.

Graph Goes Stale - Re-index After Structural Changes

There is no live file-watcher built into GitNexus yet. Re-run gitnexus analyze after significant refactors, new file additions, renamed files, or major dependency changes to keep the graph accurate. Stale graphs give Zoo Code outdated blast radius data.

Step 2 - Connect GitNexus to Zoo Code via MCP

In Zoo Code, open the MCP servers panel and click Edit Global MCP to open the global mcp_settings.json. Add the following entry inside the mcpServers object:

๐Ÿ“„ mcp_settings.json - Global MCP Config
{
  "mcpServers": {
    "gitnexus": {
      "command": "npx",
      "args": ["-y", "gitnexus@latest", "mcp"]
    }
  }
}
Fix for MCP GitNexus connection error

In case, after adding GitNexus to your MCP config, the connection showed MCP Error -32000, then bootstrap the npx cache manually:

  1. Open a terminal in your project root
  2. Run: npx gitnexus mcp
  3. Wait until you see "MCP server starting" in the output
  4. Press Ctrl+C to stop it
  5. In Zoo Code, restart MCP via the panel

If you still see MCP error -32000: Connection closed after this, also verify your Windows Firewall is not blocking Node.js - ensure an Allow (Inbound, Private) rule exists for node.exe.

Save the file. Zoo Code auto-spawns the GitNexus MCP server as a background npx process over stdio whenever a new task starts - no open ports, no cloud infrastructure, no additional configuration beyond this entry.

Project-Scoped vs. Global Config

To restrict GitNexus to one specific project instead of all workspaces, place the same JSON in .roo/mcp.json at your project root. Unlike .gitnexus/, this file can be committed to Git - it contains no secrets by default and gives your whole team the same MCP setup automatically.

Step 3 - Create the GitNexus Impact Skill

Zoo Code's Skills system loads task-specific instructions on-demand, only when your request matches the skill's description. Without a skill, you must manually tell Zoo Code to run impact analysis before every change. With this skill, Zoo Code automatically performs blast radius analysis whenever you ask it to refactor, rename, or modify anything - without cluttering the base prompt for unrelated tasks.

First, create the skill directory:

๐Ÿ–ฅ๏ธ Terminal
mkdir -p .roo/skills/gitnexus-impact

Then create the file at .roo/skills/gitnexus-impact/SKILL.md. The name field in the frontmatter must exactly match the directory name:

๐Ÿ“„ .roo/skills/gitnexus-impact/SKILL.md
---
name: gitnexus-impact
description: Analyze blast radius and dependencies using GitNexus before modifying, refactoring, or renaming any function, class, or shared utility
---

Before implementing any change to the codebase:

1. Call gitnexus_list_repos to confirm the repository is indexed and available
2. Call gitnexus_detect_changes to identify all symbols modified in the current working tree
3. Call gitnexus_context on each modified symbol to retrieve its full caller and callee graph
4. Call gitnexus_impact to get the confidence-scored blast radius of the proposed change
5. Report the risk level (LOW / MEDIUM / HIGH / CRITICAL) and all affected execution flows
6. Only proceed with implementation after presenting these findings

Risk Level Reference:
- LOW: Localized change, few or no external callers โ†’ proceed with standard testing
- MEDIUM: Several direct callers or moderate process impact โ†’ manual review of caller compatibility required
- HIGH: Numerous upstream callers across multiple critical processes โ†’ comprehensive regression testing required
- CRITICAL: Breaking change to core library or shared utilities โ†’ architectural review and phased deployment

Zoo Code discovers this skill automatically at startup by scanning the .roo/skills/ directory - no registration or configuration required. It activates only when your request semantically matches the description, keeping the base prompt clean for all other work.

Skill Location Options
  • Code mode only: Place in .roo/skills-code/gitnexus-impact/ - activates exclusively in Code mode
  • Architect mode only: Place in .roo/skills-architect/gitnexus-impact/
  • All projects globally: Place in ~/.roo/skills/gitnexus-impact/
  • Team sharing: Commit .roo/skills/ to Git - teammates get the skill automatically on clone
  • Override priority: Project .roo/ always beats global ~/.roo/; mode-specific always beats generic
Do NOT Copy Claude Code Skills

GitNexus generates agent skill files in .claude/skills/ for Claude Code's hooks system (PreToolUse / PostToolUse). These are not Zoo Code's .roo/skills/ format, so Zoo Code will not load them merely because they exist. A separate dsh compatibility plugin may consume Claude-format skills, but that is a different path and is not part of this Zoo Code setup. Always write a fresh SKILL.md using the frontmatter format above when the target is Zoo Code.

Step 3b: Generate Skills via Unity-MCP (Native Zoo Code)

As of June 2026 (PR #790), Unity-MCP includes native Zoo Code support. Zoo Code now appears directly in the AI agent dropdown inside the AI Connector window. The old workaround - selecting "Cline" and manually copying skill files from .cline/skills/ to .roo/skills/ - is no longer necessary.

Step A - Open the AI Connector panel. In the Unity Editor menu bar, click Window โ†’ AI Connector (Unity-MCP). Confirm that the panel shows Unity: Connected and MCP server: Running (http).

Step B - Select Zoo Code as the AI agent. In the AI agent dropdown, select Zoo Code. The plugin now targets Zoo Code's native configuration format - .roo/mcp.json for the MCP server config and .roo/skills/ for generated skills.

Step C - Generate the skills. In the Skills row, ensure Auto-generate is checked, then click Generate. The plugin creates all skill files directly at:

Generated output location (native Zoo Code)
flowchart TD A[".roo/"] --> B["skills/"] B --> C["<skill-name>/"] C --> D["SKILL.md"]

Zoo Code discovers all skills in .roo/skills/ automatically at startup - no copying, no directory restructuring, no registration step. Restart VS Code after generation to ensure Zoo Code picks up the new skills.

Historical Note: The Old Cline Workaround (Pre-June 2026)

Before PR #790 (merged June 1, 2026), Zoo Code was not listed in the AI agent dropdown. The documented workaround was to select Cline as the agent, generate skills into .cline/skills/, then manually copy them to .roo/skills/. If your Unity-MCP plugin lacks the Zoo Code option, update to the latest release from the GitHub releases page.

Folder Structure Must Be Preserved

Each skill must live inside its own named subfolder. The folder name must exactly match the name field in the SKILL.md frontmatter. Copying only the SKILL.md file without its parent folder will cause the skill to be silently ignored by Zoo Code.

  • โœ… Correct: .roo/skills/gitnexus-impact/SKILL.md
  • โŒ Incorrect: .roo/skills/SKILL.md
Git Tracking for Skills
  • .roo/skills/ - commit this. Skills are shareable team assets; teammates get them automatically on clone.
  • .roo/mcp.json - commit only if paths use ${env:VARIABLE} substitution (see Step 4 below). Machine-specific absolute paths will break for teammates.

Step 4 - Verify the Integration

Start a new Zoo Code task and send the following message. Zoo Code should call gitnexus_list_repos and return your indexed project name:

๐Ÿ’ฌ Verification Prompt for Zoo Code
List my indexed GitNexus repositories.

If Zoo Code returns your project name, all 7 MCP tools are live and the skill is ready. If it returns an error, ensure you have run gitnexus analyze first and that your mcp_settings.json is valid JSON (no trailing commas, no syntax errors).

The 7 GitNexus MCP Tools - Daily Usage

Once wired up, instruct Zoo Code using natural language. These are the 7 verified tools:

Ask Zoo Code toโ€ฆ Tool Invoked What It Does
"What repos are indexed?" gitnexus_list_repos Discovers all repositories registered in the global index
"Find all usages of AuthService" gitnexus_query Hybrid BM25 + semantic + RRF search; results grouped by execution process, not just file matches
"Show everything that calls parseToken" gitnexus_context 360ยฐ symbol view - all callers, all callees, imports, and definitions in one response
"What breaks if I change this function?" gitnexus_impact Full blast radius with confidence scores: LOW / MEDIUM / HIGH / CRITICAL, with affected flow names
"What does my current git diff affect?" gitnexus_detect_changes Maps changed lines in the working tree to affected execution flows before you commit
"Rename getUserById across all files" gitnexus_rename Graph-aware coordinated multi-file rename - consults the graph, not a text find-and-replace
"Run a raw graph query" gitnexus_cypher Direct Cypher queries against the underlying KuzuDB graph database for advanced traversal

GitNexus also exposes 2 guided prompts you can invoke directly: detect_impact - orchestrates detect_changes + context + impact into a single structured pre-commit risk report - and generate_map, which produces a Mermaid architecture diagram of your codebase directly from the knowledge graph.

Quick Reference - What to Commit

Commit These GitNexus Files
  • AGENTS.md - AI-readable codebase overview for all agents
  • CLAUDE.md - same overview, Claude Code format
  • .roo/skills/gitnexus-impact/SKILL.md - your impact analysis skill (shares with team)
  • .roo/mcp.json - project-scoped MCP config, if you use it instead of global
MCP Config File Locations
  • Project-local: .roo/mcp.json - inside your workspace/project folder
  • Global (all projects): mcp_settings.json - located at
    %APPDATA%\Code\User\globalStorage\rooveterinaryinc.roo-cline\settings\
Never Commit These
  • .gitnexus/ - binary KuzuDB graph database, machine-specific, auto-regenerated (already covered in your .gitignore above)

5. Live Unity Editor Access (Unity-MCP + Zoo Code)

Everything covered so far - the .gitignore, the VS Code workspace, the .rooignore, and the GitNexus knowledge graph - deals with your files. But a Unity project is more than its files. It is a live, running environment: a scene hierarchy, a serialized component graph, physics settings, materials, prefab overrides, and Play Mode state that only exists inside the Unity Editor. No amount of file reading can tell Zoo Code whether a script is attached to a GameObject, whether a prefab has broken references, or what the current scene looks like visually.

The solution is the Model Context Protocol (MCP) - an open standard that lets Zoo Code call "tools" exposed by a sidecar server process. For Unity, the right server is IvanMurzak's Unity-MCP (branded "AI Connector"). It installs entirely from a .unitypackage installer, exposes 50+ tools that operate on the live Unity Editor API, and connects to Zoo Code over a stdio pipe - no open ports exposed to the network, no cloud services, no Python or Docker required.

Why This Server Over the Others

The Unity MCP ecosystem has multiple implementations. Unity-MCP by IvanMurzak is the recommended choice for Zoo Code users because: (1) the entire Unity-side setup is a single .unitypackage - no CLI tools, no npm installs, no external runtimes required inside Unity; (2) the AI Connector window inside the Unity Editor auto-generates the exact JSON config snippet you paste into Zoo Code, so there is no guessing at paths; (3) it supports both Editor mode and Play Mode / Runtime, meaning Zoo Code can debug live game state, not just static project files; (4) the MCP server communicates over stdio - it does not open or bind any network port on your machine.

Prerequisites

  • Unity 2022.3 LTS, 2023, or Unity 6 (all confirmed supported)
  • .NET 9.0 runtime installed on your machine
  • Zoo Code VS Code extension installed and active
  • Your project already has a working .roo/mcp.json or you are using the global mcp_settings.json - both work; project-level is recommended (see the Project-Scoped Config note in Section 4)
  • Project path must not contain spaces - Unity-MCP's server process will fail to start if your project folder path includes a space character (e.g. C:/My Projects/Game will break; C:/MyProjects/Game is fine)

Step 1 - Install the Unity Plugin via .unitypackage

Go to the Unity-MCP Installation Guide on GitHub and download the latest AI-Game-Dev-Installer.unitypackage. Then, with your Unity project already open:

  1. Double-click the downloaded AI-Game-Dev-Installer.unitypackage file - Unity will open the import dialog automatically.
  2. Click Import All to bring all assets into your project.
  3. Wait for Unity to finish compiling. The AI Connector (Unity-MCP) window will be available under Window โ†’ AI Connector (Unity-MCP) once compilation is complete.
Do Not Install via CLI Inside Unity

The official installer is the .unitypackage file. Do not attempt to run npm install or any other CLI command to install the Unity-side plugin - those commands install the MCP server binary separately and are not needed here. The .unitypackage handles everything the Unity Editor needs.

Step 2 - Open the AI Connector Window

In the Unity Editor, go to Window โ†’ AI Connector (Unity-MCP). This panel is your control centre for the MCP integration. It shows the current connection status and lets you pick which MCP client to configure. Zoo Code is now a native option in the AI agent dropdown as of June 2026 (PR #790) - no Cline workaround needed. The panel generates the exact JSON snippet for your .roo/mcp.json automatically.

Internal Port - Not a Network Port

Unity-MCP uses an internal loopback connection on port 14000 between the Unity Editor plugin and the locally spawned MCP server process. This port is not exposed to the network and requires no firewall rules or port forwarding. From Zoo Code's perspective, the connection is a plain stdio pipe - Zoo Code spawns the server binary as a subprocess and communicates over standard input/output, not over TCP. You do not need to configure any port in Zoo Code's MCP settings. If port 14000 is already in use on your machine, Unity-MCP will report a binding error in the AI Connector window - free the port and restart the Editor.

Step 3 - Select Zoo Code & Generate Your Config

Inside the AI Connector (Unity-MCP) window, select Zoo Code from the AI agent dropdown and click Configure. The plugin generates a complete MCP server entry - including the "type": "stdio" transport declaration, the absolute path to the server binary, and all required connection arguments - and writes it directly to .roo/mcp.json at your project root. The generated config will look like one of the following depending on your OS:

๐Ÿ“„ Example - Windows (win-x64)
{
  "mcpServers": {
    "Unity-MCP": {
      "command": "C:/Users/YourName/Projects/YourGame/Library/mcp-server/win-x64/unity-mcp-server.exe",
      "args": [],
      "env": {}
    }
  }
}
๐Ÿ“„ Example - macOS Apple Silicon (osx-arm64)
{
  "mcpServers": {
    "Unity-MCP": {
      "command": "/Users/YourName/Projects/YourGame/Library/mcp-server/osx-arm64/unity-mcp-server",
      "args": [],
      "env": {}
    }
  }
}
Do NOT Hardcode This Path in a Shared Config

The binary path contains your local username and absolute filesystem path. Never commit a config containing it to Git - it will break for every other developer on the team. Instead, use the project-scoped .roo/mcp.json approach shown in Step 4 and add .roo/mcp.json to your .gitignore if it contains machine-specific paths. If you want to share the config, use an environment variable for the path (see the tip below Step 4).

Step 4 - Review & Customize the Generated Config

The AI Connector plugin writes the MCP server entry directly to .roo/mcp.json at your project root when you click Configure with Zoo Code selected. Open that file and add an alwaysAllow array for safe, read-only tools so Zoo Code can gather context without interrupting you with approval dialogs on every call. This is the only manual addition you need to make - everything else is auto-generated:

๐Ÿ“„ .roo/mcp.json - Unity-MCP Project Config
{
  "mcpServers": {
    "Unity-MCP": {
      "command": "/absolute/path/to/Library/mcp-server/<platform>/unity-mcp-server",
      "args": [],
      "env": {},
      "alwaysAllow": [
        "assets-find",
        "scene-get-hierarchy",
        "gameobject-get",
        "console-get-logs",
        "script-get"
      ]
    }
  }
}
Making the Config Shareable with Environment Variables

Zoo Code's mcp.json supports ${env:VARIABLE_NAME} syntax for environment variable substitution. To make the config safe to commit, each developer sets UNITY_PROJECT_PATH in their shell profile pointing to their local clone, and the shared config uses:

"command": "${env:UNITY_PROJECT_PATH}/Library/mcp-server/win-x64/unity-mcp-server.exe"

Security: What to Allow and What to Block

Unity-MCP can create, modify, and destroy assets, GameObjects, and project settings - the same power a human developer has, but executed at AI speed. The alwaysAllow list should contain only non-destructive read operations. Anything that writes, deletes, or executes code must require your explicit approval each time Zoo Code invokes it.

Tool Category Examples Policy
Read / Query assets-find, scene-get-hierarchy, gameobject-get, console-get-logs, script-get โœ… Add to alwaysAllow
Visual Capture screenshot-game-view, screenshot-scene-view โœ… Safe - add to alwaysAllow
Scene / Asset Modification gameobject-create, gameobject-component-add, assets-modify, assets-prefab-create โš ๏ธ Require approval - do NOT add to alwaysAllow
Destructive assets-delete, gameobject-destroy, package-remove ๐Ÿ›‘ Never auto-allow - always require explicit confirmation
Code Execution script-execute, reflection-method-call ๐Ÿ›‘ Never auto-allow - review every invocation carefully

Step 5 - Verify the Integration

Make sure the Unity Editor is open with your project loaded and the AI Connector window shows Connected or Connectingโ€ฆ. Then start a new Zoo Code task and send the following prompt. Zoo Code should call scene-get-hierarchy and return a live tree of your currently open scene:

๐Ÿ’ฌ Verification Prompt for Zoo Code
Use the Unity-MCP server to list all GameObjects in the current scene and tell me which ones have a Rigidbody component.

If Zoo Code returns a real scene breakdown, all tools are live. If it returns a connection error, confirm that (a) the Unity Editor is running and the project is fully loaded, (b) the AI Connector window shows Connected - if it shows a port binding error on 14000, another process is occupying that port; free it and restart the Editor, and (c) the binary path in your .roo/mcp.json matches the path shown in the AI Connector window exactly.

Unity Must Be Open

Unity-MCP communicates with the live Unity Editor process via an internal connection on port 14000 (loopback only). If Unity is closed or the project is not loaded, Zoo Code will receive a connection error from the MCP server binary. Always open the project in Unity first and confirm the AI Connector window is active, then start your Zoo Code session.

Core Unity-MCP Tools - Daily Usage

The following table covers the most commonly used tools across the 50+ available in Unity-MCP. Zoo Code selects and invokes these automatically based on your natural language instructions.

Ask Zoo Code toโ€ฆ Category What Happens in Unity
"Show me the scene hierarchy" Scene Calls scene-get-hierarchy - returns the full live GameObject tree of the currently open scene
"Find all assets named 'CarBody'" Assets Calls assets-find - queries the AssetDatabase by name, type, or path pattern
"Create a new empty GameObject called 'SpawnPoint'" Scene Calls gameobject-create - creates the object in the currently open scene
"Add a Rigidbody to the 'Player' object" Scene Calls gameobject-component-add - attaches the component via the Editor API, respecting Undo history
"Show me the Console errors" Debug Calls console-get-logs - retrieves the Unity Console output including errors, warnings, and logs
"Take a screenshot of the Game View" Visual Calls screenshot-game-view - captures what the camera currently renders, enabling Zoo Code to visually verify UI layouts or lighting
"Create a prefab from the 'Enemy' object" Assets Calls assets-prefab-create - saves the GameObject as a new Prefab asset in the AssetDatabase
"Run this C# snippet in the Editor right now" Execution Calls script-execute - compiles and runs C# code dynamically via Roslyn, essentially a REPL inside Unity

Optional - Local RAG for Large Projects

If your project has more than roughly 50 C# scripts, adding a local vector search layer prevents Zoo Code from overloading its context window when exploring the codebase. The mcp-local-rag server indexes your scripts using local embeddings and returns only the most relevant snippets when Zoo Code asks about the codebase - it does not communicate with any external service.

Add it as a second server entry in your .roo/mcp.json alongside Unity-MCP:

๐Ÿ“„ .roo/mcp.json - Unity-MCP + Local RAG
{
  "mcpServers": {
    "Unity-MCP": {
      "command": "/absolute/path/to/Library/mcp-server/<platform>/unity-mcp-server",
      "args": [],
      "env": {},
      "alwaysAllow": [
        "assets-find",
        "scene-get-hierarchy",
        "gameobject-get",
        "console-get-logs",
        "script-get"
      ]
    },
    "local-rag": {
      "command": "npx",
      "args": ["-y", "mcp-local-rag"],
      "env": {
        "RAG_HYBRID_WEIGHT": "0.6",
        "RAG_MAX_FILES": "8",
        "RAG_MAX_DISTANCE": "0.5"
      },
      "alwaysAllow": ["query_documents", "list_files"]
    }
  }
}

The recommended workflow: Zoo Code uses local-rag first to semantically locate the relevant scripts by meaning, then uses Unity-MCP to act on the live Editor objects. The RAG_HYBRID_WEIGHT value of 0.6 balances keyword and semantic search. Increase it toward 1.0 for larger codebases where exact symbol names matter more.

Production Option - Full Local RAG Stack

For projects with hundreds of C# scripts or different script types, multiple repositories, more than 10K inferrable files, or multi-user team access, neither the GitNexus or the lightweight mcp-local-rag server above may not scale. The architecture below is a fully self-hosted, production-grade RAG system built entirely on local infrastructure - no cloud dependency at any stage of embedding, retrieval, graph traversal, or generation.

Chosen Retrieval Backend: Qdrant (Self-Hosted, Local)

Qdrant is a self-hostable, open-source vector search engine written in Rust, purpose-built as a database service rather than a bare library. It is described directly by its maintainers as providing "fast and scalable vector similarity search service with convenient API," and is "tailored for extended filtering support, making it useful for all sorts of neural-network or semantic-based matching, faceted search, and other applications."

๐Ÿณ Docker - Start Qdrant locally with persistent storage
docker pull qdrant/qdrant

docker run -p 6333:6333 -p 6334:6334 \
  -v "$(pwd)/qdrant_storage:/qdrant/storage:z" \
  qdrant/qdrant

Qdrant exposes a REST API on port 6333 and a gRPC API on port 6334. Data is persisted in the mounted volume and survives container restarts.

Why Qdrant Fits Internal Software Engineering Use

Qdrant is explicitly documented and tutorialed for code retrieval. Qdrant's own engineering tutorial states its purpose plainly: to "build a semantic code search engine with Qdrant by combining natural-language and code-specific embeddings to navigate large source codebases". This is not a repurposed general-search tool, it is a validated, first-party use case.

Production precedent exists. Bloop, a production code-search tool combining semantic search, regex search, and code navigation into a single lightweight desktop application, evaluated multiple vector databases and selected Qdrant, reporting "excellent semantic search performance whilst using a reasonable amount of resources," with the ability to handle concurrent search requests while staying "fast, accurate and reliable" even on large codebases.

Independent developer confirmation supports this. Real users running local Qdrant instances for AI coding agents report: "the search speed for the LLM is incredibly fast and impressively accurate," specifically for locating elements within code. A separate documented, real-world local deployment (Qdrant plus llama.cpp plus a dedicated code embedding model) confirms this is a genuinely operable stack, not theoretical, with concrete configuration details including 3,584-dimension vectors, cosine similarity, AST-based chunking, and batched embedding requests.

Resource efficiency is proven at practical scale. One user reports: "A 1GB capacity is more than sufficient; you can embed a substantial codebase... using less than 150MB". The same source confirms Qdrant supports real-time file monitoring for incremental index updates as changes land in a repository.

Core Technical Capabilities Relevant to Multi-Client, Multi-Repo Use
  • Payload-based metadata filtering: Qdrant stores vectors alongside structured metadata and exposes a client-server API for collections, insert/upsert operations, and filtered similarity search. This allows retrieval to be scoped by repository, branch, team, service, or client, which is essential when serving multiple internal teams or multiple external clients from one shared local deployment.
  • Hybrid dense plus sparse retrieval: Qdrant supports combining dense (embedding-based) similarity with sparse (lexical/text) retrieval, addressing the well-known weakness of pure semantic search on exact identifiers such as function names, error codes, and variable names.
  • Benchmarked performance under fair, reproducible conditions: Qdrant's own benchmarking methodology explicitly states: "we do comparative benchmarks, which means we focus on relative numbers rather than absolute numbers... we run benchmarks on the same exact machines to avoid any possible hardware bias... all the benchmarks are open-sourced". Their 2024 benchmark update reports "an impressive improvement of nearly four times in certain cases" against other engines on datasets built specifically to reflect RAG at scale usage patterns.
  • Fully local and self-hostable: Qdrant runs entirely on local infrastructure with no mandatory external dependency, satisfying strict on-premise, no-internet requirements for sensitive internal or client codebases.
What Qdrant Is Not Good At

Semantic vector search has a well-documented structural blind spot for code. The relationship between a function and the things that call it is not a matter of similar wording, it lives in the call graph, the import graph, and the inheritance hierarchy. Similarity search hands back paragraphs that read as related but frequently misses the exact dependency that actually matters for a given question. This gap is what the next section addresses.

Graph-Based Code Understanding: Graphify

Qdrant and Harrier answer the question "what code is conceptually similar to this idea?" They do not answer "what calls this function?", "what breaks if I change this schema?", or "what is the shortest dependency path between the login route and the database pool?" Those are graph-traversal questions, not similarity questions, and they need a different tool: Graphify.

What Graphify Actually Is

Graphify is an open-source, actively maintained tool that turns a codebase, along with docs, PDFs, images, and video, into a queryable knowledge graph. Its own documentation is explicit about the distinction from vector search: "Not a vector index. No embeddings, no vector store: a real graph you traverse".

Instead of chunking files and searching for similar text, Graphify does the expensive reading once, up front, and compresses the project into an explicit graph of entities and the relationships between them. Answering a question afterward means walking that graph instead of re-reading files, the same way a senior engineer builds a mental map of a system once and then follows it.

How Graphify Builds the Graph

Graphify runs three separate passes over your project, and each pass has a different privacy story:

  1. Source code pass. Parsed with tree-sitter, the same category of parser a compiler uses, across roughly three dozen languages. This is entirely deterministic and local: "no LLM, nothing leaves your machine". Every relationship found this way is tagged EXTRACTED, meaning it is directly present in the code.
  2. Recordings pass. Audio or video, such as a screen-recorded design meeting, is transcribed locally with Faster Whisper. Nothing is uploaded, and the transcript folds directly into the graph.
  3. Documents pass. PDFs, whiteboard photos, and other unstructured documents cannot be parsed by a grammar, so this pass alone is sent to an AI model, either your coding assistant's model or your own configured API key. This is the only pass that can leave the machine, and it is entirely optional: running Graphify in code-only mode skips it completely, keeping the whole pipeline local.

Every edge produced by inference (as opposed to being lifted directly from source) is explicitly labeled INFERRED, and uncertain conclusions are labeled AMBIGUOUS. This confidence tagging is a deliberate design choice that prevents the graph from presenting a guess with the same authority as a verified fact.

Once built, Graphify runs Leiden community detection over the graph, so modules fall out of the actual shape of the code: authentication clusters separately from billing, infrastructure clusters on its own, with the bridge files between them visibly highlighted, without anyone having drawn that map by hand.

Published Performance Data

Graphify's maintainers publish a benchmark methodology alongside their numbers rather than a bare headline figure. Their disclosed results:

Benchmark Metric Graphify Comparison field
LOCOMO (n=300) Recall@10 0.497 mem0: 0.048, supermemory: 0.149
LOCOMO (n=300) QA accuracy 45.3% supermemory: 49.7%, mem0: 27.3%
LongMemEval-S (n=50) QA accuracy 76% tied with dense RAG
Graph construction LLM credits used 0 per-token cost for most comparable systems

They describe their validation approach as: "every system ran on the same harness with the same model and budgets, scored by a judge blind-validated against a second judge (90.6% agreement, Cohen's kappa 0.81)". This is a more transparent methodology than a bare marketing claim, but it remains self-reported and has not been independently reproduced. On raw question-answering accuracy, Graphify does not universally win, supermemory scores higher on LOCOMO, and it only ties dense RAG on LongMemEval-S. Its decisive advantages are recall and zero ongoing LLM cost to build and maintain the graph, not superior answer accuracy in every case.

Why Graphify Complements Qdrant Instead of Replacing It
Question type Best tool Why
"What calls this function, and what does it call?" Graphify Structural, exact, graph traversal
"What breaks if I change this database schema?" Graphify Dependency path, not semantic similarity
"Which client repositories reference a similar authentication pattern?" Qdrant Cross-repo, conceptual similarity, needs metadata filtering
"Show me code similar to this snippet, even if unrelated structurally" Qdrant Pure semantic similarity
"Find the shortest dependency path between two modules" Graphify Native path-finding operation
"Search company knowledge base for design docs on this topic" Qdrant Unstructured document retrieval across teams and clients

Graphify is fundamentally scoped to a single repository or corpus, its output folder sits directly next to the code it analyzed. It has no built-in multi-tenant access control, no cross-repository query layer, and no client-isolation model. Qdrant was specifically chosen earlier in this architecture for exactly those capabilities: payload-based filtering by repository, team, branch, and client. Running both together closes the gap neither one covers alone: Graphify handles exact structural questions within a codebase, Qdrant handles conceptual search across many codebases, clients, and document types.

Step-by-Step: Installing and Using Graphify Alongside Qdrant

This walkthrough assumes the Qdrant, Harrier, and Qwen 3.6 stack from the sections above is already running locally.

Step 1: Install Graphify

Graphify ships as a package installable through your coding assistant's plugin or skill mechanism, and as a standalone CLI. Install it into the project you want mapped, following the quickstart flow documented by the project: "Install graphify, map your repo into a knowledge graph with /graphify, query it from the CLI, wire up the MCP server, and review PRs against the graph, all in about five minutes".

๐Ÿ–ฅ๏ธ Terminal - Install Graphify
# Standalone CLI install
npm install -g graphify
# or, inside a supported coding assistant (Claude Code, Cursor, Codex, Gemini CLI, Windsurf, etc.)
# the /graphify slash command is available once the extension/skill is added
Step 2: Build the Graph

From inside your coding assistant, or from the CLI, point Graphify at the project directory:

๐Ÿ–ฅ๏ธ Terminal - Build graph
graphify build ./my-repo

Or, inside Claude Code, Cursor, or Codex:

๐Ÿ’ฌ Coding assistant
/graphify ./my-repo

This runs the three passes described above (code, recordings, documents) and writes the full graph to a graphify-out/ folder placed directly next to your code. No files are modified; the raw source remains untouched, matching the same "never edit the source" discipline used in the Karpathy Wiki pattern this tool is philosophically related to.

To keep everything fully local, including skipping the one pass that can leave the machine, use code-only mode:

๐Ÿ–ฅ๏ธ Terminal - Code-only mode
graphify build ./my-repo --code-only
Step 3: Enable Incremental Rebuilds

Graphify hashes every file with SHA-256, so subsequent runs only re-read what changed, rather than rebuilding the entire graph from scratch. Wire this into your commit workflow so the graph never goes stale:

๐Ÿ–ฅ๏ธ Terminal - Install git hook
graphify hook install

This installs a git hook that rebuilds the graph automatically after every commit, using the AST-only pass, at no ongoing LLM cost. Graphify also installs a git merge driver so the underlying graph.json file does not get left with unresolved conflict markers when two engineers commit in parallel on different branches.

Step 4: Expose the Graph as an MCP Server

To let any AI coding assistant query the graph as a native tool rather than reading raw files, start the MCP server:

๐Ÿ–ฅ๏ธ Terminal - Start MCP server
graphify mcp start

For a team setting where multiple engineers need to query the same graph concurrently, run the shared HTTP mode instead of a local-only MCP instance, so the whole team queries one graph rather than maintaining separate local copies:

๐Ÿ–ฅ๏ธ Terminal - Shared HTTP mode
graphify mcp start --http --port 8765

Register this MCP endpoint with each engineer's coding assistant configuration, following the same connection pattern used for any other MCP tool.

Step 5: Query the Graph

Once running, ask structural questions either directly through your assistant (which will call the MCP tool automatically) or via the CLI:

๐Ÿ–ฅ๏ธ Terminal - Query the graph
graphify query "what calls handleStripeWebhook"
graphify path "auth-service" "database-pool"
graphify explain "billing-service"
  • query walks the graph for a direct answer instead of opening files.
  • path finds the shortest dependency chain between two named entities, useful for "what breaks if I change this" questions.
  • explain gives a plain-English tour of any entity, function, module, or service, built from its position in the graph.
Step 6: Run Both Systems Side by Side

With both systems running, route queries by type rather than picking one exclusively:

Query routing diagram
flowchart TD Q["Developer question"] --> S["Structural / dependency / &quot;what calls this&quot; question"] S --> G["Graphify MCP server"] Q --> C["Conceptual / cross-repo / &quot;find something similar&quot; question"] C --> QD["Qdrant (via Harrier embeddings)"] Q --> L["Local LLM (Qwen 3.6 Dense 27B for coding, MoE 35B for general use)"]

Most coding assistants can be configured to call both MCP tools and let the model decide which one is relevant to a given question, or you can explicitly instruct the assistant in its system prompt to prefer Graphify for "who calls this" and "what depends on this" style questions, and Qdrant for "find something like this" or "search across all our client repositories" style questions.

Step 7: Maintain the Graph

Because code parsing is deterministic and AST-based, Graphify's structural facts (EXTRACTED edges) do not decay the way LLM-inferred knowledge can. The only ongoing maintenance is confirming the git hook remains installed after major repository restructuring, and periodically spot-checking INFERRED and AMBIGUOUS edges if the documents pass is enabled, since that is the one pass that involves model reasoning rather than pure parsing.

Chosen Embedding Model: Harrier-oss-v1-0.6B

Qdrant stores and searches vectors, but it does not generate them, an embedding model is required to convert code and documents into vectors before indexing. Harrier-oss-v1-0.6B was chosen because embedding quality is the priority for this system, and it is the objectively stronger model on every published multilingual benchmark evaluated.

A 2026 open-source embedding benchmark across 14 models found: "AVG nDCG@3 = 0.8911... the highest-ranked open-source row that ships with a license suitable for commercial use without restrictions". The same benchmark directly compared it against the closest architecturally comparable alternative: "Harrier-0.6b (0.8911) sits 0.074 nDCG@3 above Qwen3-Embedding-0.6B (0.8168), built on the identical Qwen3-0.6B base. The training corpus and instruction recipe drove the gap, not the parameter count". Harrier ships under an MIT license, making it fully suitable for internal commercial deployment without restriction.

Chosen Local LLMs

Use case Model
Coding and software engineering tasks Qwen 3.6, Dense, 27B
General / other use cases Qwen 3.6, MoE, 35B

Both models run fully locally alongside the Qdrant retrieval layer, the Graphify knowledge graph, and the Harrier embedding model, ensuring no code, client data, or proprietary information leaves the internal network at any stage of embedding, retrieval, graph traversal, or generation.

Multi-User Access: Tailscale

Everything above runs on a single local machine or internal server, but multiple engineers, teams, or clients need secure access to that same Qdrant instance, Graphify MCP server, and LLM endpoints without exposing them to the public internet. Tailscale is the recommended networking layer for this.

Tailscale is described as "a secure networking solution that streamlines connecting devices and services securely across different networks. It enables encrypted point-to-point connections using the open source WireGuard protocol, which means only devices on your private network can communicate with each other. Tailscale creates a peer-to-peer mesh network (known as a tailnet)". This directly matches the requirement for this architecture: connecting devices across different locations directly (peer-to-peer) via a private mesh network, without port forwarding, complex firewall rules, or dynamic DNS.

Why It Fits This Architecture
  • No port forwarding or firewall configuration required. Tailscale's own documentation confirms: "connections between tailnet devices work seamlessly across firewalls and Network Address Translation (NAT) without requiring port forwarding or complex firewall rules". This is critical for a locked-down internal server hosting Qdrant, the Graphify MCP server, and the LLM endpoints, since no inbound ports need to be opened to the internet.
  • End-to-end encryption by default. Tailscale is "built on top of WireGuard, a modern VPN that provides end-to-end encryption between devices. Tailscale cannot read your traffic". Combined with WireGuard, Tailscale "constructs a mesh network topology with additional network services and authentication mechanisms," while WireGuard alone only provides encrypted tunnels between two endpoints.
  • Peer-to-peer by design, avoiding centralized bottlenecks. "The Tailscale approach avoids centralization where possible, resulting in both higher throughput and lower latency as network traffic can flow directly between devices".
  • Identity-based access control for multiple users. Tailscale is positioned as "a Zero Trust identity-based connectivity platform... connects remote teams, multi-cloud environments, CI/CD pipelines, Edge and IoT devices, and AI workloads". It supports "RBAC policies to determine which users, roles, or groups can access, which nodes on your tailnet," and "automates user and group provisioning to safeguard against unauthorized resource access via SCIM integrations with leading identity providers".
  • Fast to deploy, no networking expertise required. "You can deploy a tailnet in minutes without requiring extensive configuration, server setup, and networking expertise".
  • Directly validated for local AI stacks. A documented walkthrough builds an equivalent architecture, a fully offline AI lab running local LLMs via Ollama, made securely accessible through Tailscale: "with Tailscale Serve, you can access it securely from anywhere on your Tailnet... complete with TLS and no reverse proxy configuration required... you'll have a system that can run LLMs, fully offline".
How Many Users Can Use the System

Tailscale's mesh model scales by adding authenticated devices (users) to the same tailnet rather than by provisioning new public endpoints per user. Each additional engineer, internal team member, or client is enrolled as a new node on the tailnet, inheriting the access policy assigned to their role or group. The practical user limit is governed by your Tailscale plan and node limits and access-control policy design, not by networking complexity.

Complete Local Architecture

Local RAG Architecture - Full Stack
flowchart TD U["Authorized users (internal teams / clients)"] --> T["Tailscale (WireGuard-based private mesh network)<br/>- No port forwarding, no exposed public endpoints<br/>- Identity-based, role/group access control<br/>- Encrypted peer-to-peer connections"] T --> G["Graphify MCP server<br/>- Structural code graph<br/>- Call graphs, dependency paths, module clusters<br/>- EXTRACTED / INFERRED confidence labels"] T --> H["Harrier-oss-v1-0.6B<br/>(embedding model)"] H --> Q["Qdrant (self-hosted, local)<br/>- Cross-repo conceptual search<br/>- Multi-client metadata filtering<br/>- Docs, ADRs, schemas, runbooks"] G --> L["Local LLM inference<br/>- Qwen 3.6 Dense 27B -> coding tasks<br/>- Qwen 3.6 MoE 35B -> general tasks"] Q --> L

Metadata attached to every Qdrant-indexed chunk (repository, branch, commit, file path, symbol, language, service, team, client, access scope) allows conceptual retrieval to be correctly scoped per internal team or per client. Graphify's confidence-labeled edges (EXTRACTED, INFERRED, AMBIGUOUS) let structural queries be trusted at the right level of certainty. Tailscale's access-control layer ensures only authorized devices can reach either system in the first place.

Alternatives Considered and Why They Were Not Chosen

FAISS was not chosen as the retrieval backend because it is a similarity-search library, not a database service. It provides no native multi-user concurrency handling, no built-in metadata filtering, no persistence layer, and no CRUD support, all of which have to be custom-built around it to serve multiple users safely. Wrapping it (via community projects or Meta's own reference RPC script) shifts significant engineering and operational risk onto the implementing team, with no independently verified production benchmark proving it outperforms a purpose-built engine like Qdrant for this use case.

Semble was not chosen as the primary or required component, because its documented strength, extremely fast, token-efficient, single-repository code search using a lexical plus semantic hybrid technique, solves a narrower problem (minimizing per-agent token spend on the active repository) rather than the broader requirement here: a shared, multi-user, multi-repository, access-controlled retrieval system spanning code, documentation, schemas, and internal knowledge across teams and clients.

Qwen3-Embedding-0.6B was not chosen as the primary embedder, but remains a good alternative in specific circumstances: when adjustable vector size is needed for storage and index efficiency, or when more flexibility is required to host the embedder privately on third-party hosting providers as a ready-to-use service in the future. Both models are open-weight and commercially usable (Harrier: MIT; Qwen3: Apache-2.0), so licensing does not favor either one, the choice comes down to raw embedding quality (Harrier) versus deployment flexibility (Qwen3).

Graphify was not chosen as a replacement for Qdrant, because it solves a different, narrower problem. It has no cross-repository query layer, no client or tenant isolation, and its own published benchmarks show it does not universally beat dense RAG on question-answering accuracy, only on recall and zero-cost graph construction. It was added alongside Qdrant specifically because it closes a real, well-documented gap that vector similarity search cannot close on its own: exact structural relationships between code entities.

Project Structure Best Practices for AI-Assisted Unity

The quality of Zoo Code's Unity-MCP session is directly affected by how your project is organized. Disorganized hierarchies and inconsistent naming are the primary causes of agentic failure - the AI wastes tokens searching for things that would be obvious in a well-structured project.

  • CamelCase everything. Unity's command-line tools and MCP path parsers break on spaces in file or folder names. Use Assets/Materials/CarBody_Red.mat, not Assets/Materials/car body red.mat.
  • Keep .meta files in Zoo Code's visibility range. Our .rooignore blocks *.meta files from AI indexing - that is correct for file search. However, some Unity-MCP tools (like assets-find and find-references) rely on the GUIDs inside .meta files to track asset dependencies. The MCP server reads them directly from disk, bypassing .rooignore, so they must always be present and committed.
  • Use C# Namespaces. Namespaces allow Unity-MCP's reflection tools to separate modules cleanly, preventing naming conflicts and making the dependency graph much easier for Zoo Code to navigate. A file without a namespace forces the AI to scan broader context to understand where it belongs.
  • Maintain a consistent folder contract. If all materials live in Assets/Materials/, Zoo Code can form a reliable mental model of the project and will waste fewer tool calls on redundant discovery searches.
  • Keep a project context file. A CLAUDE.md or AGENTS.md at the project root (auto-generated by GitNexus in Section 4, or written manually) gives Zoo Code a high-level architecture overview at the start of every session. Include naming conventions, key systems, and any Unity version specifics. This file dramatically reduces the number of exploratory tool calls Zoo Code needs to orient itself.
Quick Reference - Unity-MCP File Handling
  • .roo/mcp.json - safe to commit if paths use ${env:VARIABLE} substitution instead of absolute paths
  • CLAUDE.md / AGENTS.md - your project context file; teammates and AI agents both benefit from it
  • Library/mcp-server/ - the compiled server binary is machine-specific and auto-regenerated by the Unity plugin; it is already excluded by the Library/ rule in your .gitignore above
  • .roo/mcp.json containing hardcoded absolute paths - this will break every other developer's setup silently; use ${env:UNITY_PROJECT_PATH} for team sharing

5b Grok + AI Game Developer MCP

This section is an additional setup path for Grok. It does not replace the Zoo Code Unity-MCP guidance in Section 5. The verified Grok connection uses the AI Game Developer MCP server running locally from the Unity project, while Grok connects to that server over HTTP.

Verified Architecture

The verified connection chain is Unity Editor โ†’ AI Game Developer MCP package โ†’ local gamedev-mcp-server โ†’ localhost HTTP endpoint โ†’ project .grok/config.toml โ†’ Grok session. The active MCP server name is ai-game-developer. The verified source project used http://localhost:20085; that number is not a universal port.

Use the new project's endpoint

The package derives its default port from the project directory path. Read the live localhost URL from the Unity MCP panel for each new project, then put that exact URL in Grok's configuration. Do not copy 20085 unless the new project's panel also reports it.

The verified project used Unity 6000.0.78f1 with com.ivanmurzak.unity.mcp version 0.83.1. The local process listened only on localhost; do not expose it through a public URL, LAN address, tunnel, reverse proxy, or port forwarding.

Project Configuration

First open the new Unity project, install and start the AI Game Developer MCP package, and read the displayed local server URL in its Unity panel. At the new project root, create .grok/config.toml with the following configuration. Replace PORT with the port actually shown by Unity.

.grok/config.toml
[mcp_servers.ai-game-developer]
url = "http://localhost:PORT"
enabled = true

Open a new Grok session with that Unity project as its workspace. If the configuration was added or changed while Grok was already open, restart the session or reload MCP so it reads the project-local configuration again. A user-level Grok configuration can point to a different Unity project, so the project-local file is the safer source of truth for multi-project work.

You may independently confirm the endpoint after reading the port from Unity:

Windows PowerShell
netstat -ano | findstr :PORT

The expected listener is the local gamedev-mcp-server process started for the current Unity project. A same-folder project normally receives the same default port; a different project directory normally receives another port in the package's 20000โ€“29999 range.

Safe Connection Check

Use a read-only health check before asking Grok to edit scripts, scenes, prefabs, packages, project settings, or Play Mode. This confirms that the session reaches the intended Unity Editor without granting write authority.

First safe Grok prompt
Using only the MCP server ai-game-developer, do not modify anything. Run tool discovery or unity-tool-list, then scene-list-opened and console-get-logs with a small maxEntries value. Report the server name, whether tools respond, and the opened scene path.

In the verified project, unity-tool-list, scene-list-opened, and console-get-logs responded successfully. This check establishes connectivity only; it does not authorize state changes.

Tool category Examples Session policy
Read-oriented unity-tool-list, scene-list-opened, console-get-logs, assets-find Use first to verify the target editor and inspect state
State-changing assets-modify, object-modify, scene-save, tests-run Allow only when the task expressly requests the action
Disruptive Script execution, deletes, package changes, Play Mode controls Require explicit scope and a clear expected outcome

Troubleshooting

Symptom Likely cause Verified corrective action
Grok cannot connect or no Unity tools appear Unity is closed, the package server is not running, or the configured port differs from the Unity panel Open the target project, start its MCP server, copy the panel URL into .grok/config.toml, then reload MCP or start a new session
Grok targets an old or wrong Unity project A user-level config still points to another project's port Use the current project's .grok/config.toml, confirm Grok opened that workspace, and check the endpoint against the active Unity panel
A tool appears in Unity's list but Grok says Tool not found The Grok bridge may expose a subset of Unity's advertised tools in that session Use tools proven by successful dispatch, beginning with the health-check tools; never claim an unavailable tool ran
Tools fail after scripts change or Unity compiles Domain reload or MCP server restart is in progress Wait for Unity to finish compiling, confirm the MCP panel is healthy, then retry a read-only check

Grok Setup Checklist

  • Unity is open, compiled, and running the AI Game Developer MCP server
  • The live localhost endpoint was copied from the Unity MCP panel
  • The project contains .grok/config.toml with server name ai-game-developer
  • The endpoint in the TOML file matches the active Unity project, not a previous project
  • A new Grok session was opened or MCP was reloaded after configuration
  • unity-tool-list, scene-list-opened, and console-get-logs respond before any modification task begins

5c Grok CLI Sandboxing and Permissions

This additional section covers containment for the Grok Build CLI itself. It is separate from Unity MCP connectivity: MCP determines how Grok reaches the Unity Editor, while the sandbox and permissions determine where Grok and its child processes may operate and which tool requests the model may make.

Start in the Project Directory

Start Grok in the project directory it should work in, or pass that directory explicitly with --cwd. The built-in strict sandbox permits reads from the current working directory and necessary system paths, permits writes only in the current working directory, /tmp, and ~/.grok/, and blocks child-process network connections on Linux.

Shell
cd /path/to/project
grok --sandbox strict

Or use an explicit working directory:

Shell
grok --cwd /path/to/project --sandbox strict
Strong containment, not complete isolation

Strict mode is applied when Grok starts and cannot be removed during that process. It still needs access to system paths and ~/.grok/ for Grok state. Child-process network blocking is enforced on Linux; the documented macOS sandbox does not currently enforce this network block. Do not present strict mode as a complete OS or VM isolation boundary.

Permission Guardrails

Do not use --always-approve or --yolo for normal project work. Permissions decide which tool actions the model may request; the sandbox independently constrains what an approved process can actually access. Deny rules take precedence over allow rules.

For non-interactive or headless work, dontAsk denies requests that are not explicitly allowed. The example below allows project reads, searches, file edits, and Git commands only. Add MCP permissions separately if you intend Grok to call Unity MCP tools.

Headless guarded example
grok --cwd /path/to/project \
  --sandbox strict \
  --permission-mode dontAsk \
  --allow 'Read' \
  --allow 'Grep' \
  --allow 'Edit(**/*)' \
  --allow 'Bash(git *)'

Use this pattern only when unattended edits are genuinely intended. For an interactive session, the documented default permission mode is ask, which prompts before actions requiring approval. Keep the allow list narrow and add explicit denies for operations you never want the agent to request.

Example deny rule
--deny 'Bash(rm -rf *)'

A sandbox does not automatically authorize Unity changes. If the Unity MCP is configured, use read-only Unity MCP checks first and explicitly scope any scene, asset, script, package, test, or Play Mode action in the prompt and applicable permission rules.

Persistent Project Sandbox Profile

For a reusable project policy, create .grok/sandbox.toml in the project root. Custom profiles may extend a built-in profile such as strict. The deny list is enforced before any permissions and supports paths and glob patterns.

.grok/sandbox.toml
[profiles.project-only]
extends = "strict"
restrict_network = true
deny = ["/path/to/private", "**/.env", "**/*.pem"]

Launch using the profile name:

Shell
grok --cwd /path/to/project --sandbox project-only

Use a real absolute private-path value in deny; do not leave the placeholder in a policy you depend on. A project .grok/sandbox.toml defines the profile, while activation can occur through --sandbox, GROK_SANDBOX, or a sandbox profile setting in configuration. Built-in names such as strict cannot be redefined as custom profiles.

Sandbox Checklist

  • Start Grok with the intended Unity project as --cwd
  • Use --sandbox strict for untrusted repositories or tighter local containment
  • Use the default interactive ask mode unless unattended automation is required
  • For unattended execution, use dontAsk and narrow explicit allow rules
  • Never enable unrestricted approval merely to avoid prompts
  • Use a project .grok/sandbox.toml when the same restrictions should travel with the project
  • Verify the operating system's enforcement limits, especially child-process network restrictions on macOS

6 GitHub Copilot Chat (alongside Zoo Code)

As of 2026, GitHub Copilot Chat has matured into a full multi-agent, MCP-aware platform with agent plugins, plan memory, fork/checkpoint workflows, and its own MCP marketplace. It does not replace Zoo Code - but it fills a real gap: inline completions, quick refactors, and GitHub-centric workflows. This section shows you how to set it up safely and make it coexist cleanly with your existing Zoo Code + GitNexus stack.

Read the Privacy Section First

As of April 24, 2026, GitHub uses Copilot interaction data from Free, Pro, and Pro+ users to train AI models by default. This includes your prompts, accepted suggestions, code snippets, cursor context, and file names. You must opt out before using Copilot for any proprietary code. Steps are in the Privacy section directly below.

Step 1 - Install & Sign In

  1. In VS Code Extensions, install GitHub Copilot and GitHub Copilot Chat (often bundled with recent VS Code builds - check the Extensions panel for updates).
  2. After install, click the Copilot icon in the status bar โ†’ Use AI Features with Copilot for Freeโ€ฆ โ†’ sign in with your GitHub account.
  3. If you use both a personal and a work GitHub account, assign the correct one per workspace via: Accounts menu (bottom-left avatar) โ†’ Manage Extension Account Preferences โ†’ GitHub Copilot Chat.

Step 2 - CRITICAL: Stop GitHub Training on Your Code

Copilot Business and Enterprise plans are exempt - their data is never used for model training and the toggle does not appear in their settings. If you are on Free, Pro, or Pro+, you must opt out manually.

Do This Now - Takes 10 Seconds
  1. Go to: https://github.com/settings/copilot/features
  2. Find the option: "Allow GitHub to use my data for AI model training"
  3. Toggle it OFF / Disabled
  4. Save.

This does not stop GitHub from processing your code to generate suggestions - that is inherent to how Copilot works. It only stops GitHub from using your interaction stream as training data for future models. The opt-out applies retroactively to previously collected data as well.

Step 3 - Disable VS Code Telemetry & Feedback Signals

Add both of these to your user-level settings.json (Ctrl+Shift+P โ†’ "Open User Settings (JSON)"):

๐Ÿ“„ settings.json (user-level - applies to all workspaces)
{
  "telemetry.telemetryLevel": "off",
  "telemetry.disableFeedback": true
}
  • telemetry.telemetryLevel: "off" - silences all VS Code telemetry events (usage stats, crash reports, feature flags).
  • telemetry.disableFeedback: true - removes the ๐Ÿ‘ ๐Ÿ‘Ž feedback buttons and the issue reporter from Copilot Chat. These buttons send interaction data back to Microsoft as training signals. After restarting VS Code the thumbs disappear from the Copilot Chat panel entirely.

Step 4 - Privacy-Friendly Companion Extensions

A. DeepSeek V4 for Copilot Chat (Cloud API)

Install Vizards.deepseek-v4-for-copilot from the VS Code Marketplace. This adds DeepSeek V4 Pro and DeepSeek V4 Flash as selectable models inside the Copilot Chat model picker. Agent mode, MCP tool calls, custom instructions, and VS Code tools all continue to work - Copilot routes the request to DeepSeek's inference API instead of GitHub-hosted models. Requires a DeepSeek API key from platform.deepseek.com.

Known Bug & Data Jurisdiction Caveat

There is an active bug where multi-turn conversations with DeepSeek reasoning models throw 400 errors because Copilot does not correctly round-trip reasoning_content fields in conversation history. Single-turn calls and non-reasoning model variants work fine. Check the extension's GitHub issues for current fix status before relying on it for long agentic sessions.

Your prompts and code context go to DeepSeek's servers via your API key - not GitHub's. DeepSeek is a Chinese provider subject to Chinese data law. Do not use this extension for NDA-covered code, unreleased IP, or regulated personal data (GDPR, HIPAA).

B. DeepSeek via Ollama - Fully Local (Zero Cloud Exposure)

Search the VS Code Marketplace for "DeepSeek for GitHub Copilot" (the Ollama-backed extension). This runs a DeepSeek model entirely on your local machine - zero data leaves your machine. Install Ollama, start it, install the extension, and use @deepseek as the participant inside Copilot Chat. On first invocation it auto-downloads the selected model. This is the correct choice for proprietary Unity code where you want the Copilot Chat UX but absolutely no cloud exposure.

C. Chat Customizations Evaluations (Microsoft)

Available from microsoft/vscode-chat-customizations-evaluation on GitHub. This extension analyzes your Copilot instruction files, prompt files, and skill files locally inside VS Code - it identifies ambiguities, overly dense instructions, and potential accidental secret leakage in your instruction context. Because Copilot includes your instruction files in every chat request, keeping them clean and free of sensitive strings is a meaningful privacy action.

Step 5 - Configure Agent Mode

  1. Open Settings and search github.copilot.chat.agent, or use Command Palette โ†’ "GitHub Copilot: Configure Agent Mode".
  2. Ensure agent mode is enabled (it is on by default in recent VS Code builds).
  3. Keep auto-approve / yolo mode OFF by default. Only enable it for throwaway branches or fully sandboxed tasks with no production assets at risk.
  4. Run /init inside a new Copilot Chat session on a new project - Copilot will analyze your codebase and generate a project-specific instruction context automatically.

Store project-level instructions in .github/copilot-instructions.md at your repo root. Copilot loads this file as a persistent system prompt for all chat and inline suggestion contexts, shared automatically across the team on clone. Run the Chat Customizations Evaluations extension (Step 4C) on this file periodically to keep it focused and safe.

Step 6 - Disable AI Features Per Sensitive Workspace

For repositories that must never send code to any external model - client work under NDA, unreleased game content, regulated data - add the following to .vscode/settings.json inside that specific repo:

๐Ÿ“„ .vscode/settings.json (per sensitive repo - commit this file)
{
  "chat.disableAIFeatures": true
}

This disables inline completions and Copilot Chat for that workspace only. Copilot remains fully active in all other projects. Commit this file so every teammate opening this repo gets the same protection automatically.

Recommended VS Code Profile Setup

VS Code Profiles let you maintain two completely separate extension environments with one click:

  • Profile: "AI-Assisted" - Copilot + DeepSeek extensions + Zoo Code + all MCP servers active. Use for normal, non-sensitive development work.
  • Profile: "Local-Only" - Zoo Code + local Ollama model only. Copilot extension fully disabled. Open sensitive client repos, NDA work, or regulated data exclusively in this profile.

Switch profiles via the gear icon (bottom-left) โ†’ Profiles โ†’ Switch Profile. Each profile maintains its own independent extensions list, settings overrides, and keybindings.

Copilot Chat + Zoo Code - Division of Labor

Running both tools simultaneously is the recommended setup for Unity developers. They do not conflict - each targets a different workflow layer. The key is knowing which tool to reach for.

Workflow Best Tool Reason
Inline ghost-text completions while typing โœ… Copilot Native to Copilot; Zoo Code does not provide inline ghost text
Quick explain / refactor / test gen from a selection โœ… Copilot Right-click โ†’ Copilot context menu is faster for single-file bounded tasks
GitHub PR review, issues, Copilot Coding Agent โœ… Copilot Copilot has native GitHub integration; Zoo Code has none
GitNexus graph-aware blast radius analysis ๐Ÿ”ด Zoo Code GitNexus MCP is configured only in Zoo Code (see Section 4)
Live Unity Editor access via Unity-MCP ๐Ÿ”ด Zoo Code Unity-MCP is configured only in Zoo Code's MCP stack (see Section 5)
Multi-step agentic refactors with per-tool approval ๐Ÿ”ด Zoo Code Zoo Code's per-tool approval granularity and mode system gives finer control
Sensitive / NDA / proprietary code (no cloud) ๐Ÿ”ด Zoo Code + local model Route to a local Ollama model in Zoo Code; disable Copilot for that workspace (Step 6 above)
Custom model switching (Claude, Gemini, GPT, Ollama) Either Both tools support custom model selection in their respective settings panels

7 How to Use Local AI Models for Development (Qwen3.8-27B × Unity 6 - fully offline)

In plain terms: this section explains how to run an AI coding assistant entirely on your own PC - no cloud, no subscriptions, no data leaving your machine. A Qwen3.8-27B model (the "brain") is served by llama.cpp, driven by the DeepSeek Harness (dsh) (the "assistant app"), and connected to the Unity 6 Editor through Unity-MCP so it can inspect scenes, assets, and code for you. It can also search your code by meaning with a local RAG sidecar, navigate code with compiler-grade accuracy through a language server, fetch web documentation on demand, and pause for your approval before every scene or asset write.

  • Hardware: a high end PC with a 90-class RTX GPU (24GB VRAM), a 9-series Ryzen CPU, and 32GB of RAM.
  • Install first: Node.js 24 or newer (for dsh) ยท Docker Desktop (for Qdrant) ยท .NET 10 SDK (for the C# language server) ยท Python 3.12 (for the RAG sidecar) ยท Git.

Architecture Overview

The finished stack has two layers. Layer 1 is the model itself: llama.cpp serves Qwen3.8-27B through an OpenAI-compatible API at http://127.0.0.1:8080/v1, reaching about 51 tokens per second with MTP speculative decoding. Layer 2 is the agent harness: DeepSeek Harness (dsh) gives you a web UI at http://127.0.0.1:3080, three presets, context compaction at 85 percent of the 80K window, human approval gates on Unity writes, and a skills layer for engineering discipline. Around the harness sit four helpers: Unity MCP for editor control, a fetch MCP for web pages, a C# language server for exact code navigation, and a RAG sidecar for semantic code search.

๐Ÿ—๏ธ System architecture
flowchart TD subgraph L1["Layer 1: LLM Inference (llama.cpp)"] A["Qwen3.8-27B in UD-Q4_K_XL quantization
OpenAI API at 127.0.0.1:8080/v1
~51 tokens/s with MTP speculative decoding"] end subgraph L2["Layer 2: Agent Harness (dsh)"] B["Web UI at 127.0.0.1:3080
Presets: Code Edit / Deep Debug / Mechanical
Compaction at 85% of 80K window
HITL approval gates on Unity writes
Skills: grill-with-docs, diagnosing-bugs, tdd, handoff"] end A -- "OpenAI-compatible API" --> B B --> U["Unity MCP
scenes, GameObjects, tests, console, C# execution"] B --> F["Fetch MCP
web page retrieval"] B --> S["LSP (csharp-ls)
exact code navigation"] B --> R["RAG sidecar
Qdrant (Docker) + Harrier embedder + bge-reranker-v2-m3"]

The Language Model (llama.cpp)

llama.cpp is a C++ program that runs AI models locally. It exposes an API that looks exactly like OpenAI's, so any tool built for cloud models can talk to it instead, while the data never leaves your machine.

The model chosen is Qwen3.8-27B, a 27 billion parameter model. It is downloaded in a compressed format called a quantization, specifically UD-Q4_K_XL, which shrinks the model from about 55GB down to about 18GB with almost no quality loss.

Setup

  1. Download the prebuilt llama.cpp Windows + CUDA binary and extract it to C:\llama-cpp\.
  2. Download the model from HuggingFace (unsloth/Qwen3.8-27B-GGUF), specifically the file Qwen3.8-27B-UD-Q4_K_XL.gguf (17.9GB).
  3. Verify the download by comparing its SHA-256 hash against the hash published in the repository.
  4. Launch the server with the command below.
๐Ÿ–ฅ๏ธ Server launch (PowerShell)
& "C:\llama-cpp\llama-server.exe" `
  -m "C:\Qwen3.8-27B-UD-Q4_K_XL.gguf" `
  --fit on -fitt 512 -c 81920 -fa on --jinja `
  --cache-type-k q8_0 --cache-type-v q8_0 `
  --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.7 `
  --cache-reuse 256 -t 16 --parallel 1 `
  --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0

What the important flags mean, in plain language:

  • --fit on -fitt 512: let the server automatically decide how much of the model fits in GPU memory, reserving 512MB as a safety buffer. Never pin the GPU layer count manually; doing so caused an out of memory crash during setup.
  • -c 81920: set the context window to 80,000 tokens (about 60,000 English words). This is how much text the model can hold in mind during a single conversation.
  • -fa on: enable flash attention, a faster way to compute attention that saves memory.
  • --cache-type-k q8_0 --cache-type-v q8_0: store the model's memory cache at 8-bit precision, halving VRAM usage with negligible quality loss.
  • --spec-type draft-mtp: enable Multi-Token Prediction. The model guesses several tokens at once, which speeds up output by roughly 50 percent. The measured acceptance rate was 74 to 78 percent.
  • --temp 1.0 --top-p 0.95: sampling parameters. Temperature 1.0, with thinking enabled, measured 18.7 percentage points better on coding accuracy than temperature 0.2.

Measured performance

Three configurations were benchmarked. The governing law discovered: decode speed scales with how many model weights spill to CPU. The 80K context configuration with a 512MB fit buffer is dramatically faster because only 1.1GB sits on CPU, versus 3GB or more at higher contexts.

Configuration Context GPU Layers CPU Spill Decode Speed
Q4_K_XL at 98K, fitt 2048 98K 55/66 3.07 GiB 26 to 28 tok/s
Q6_K at 131K, fitt 2048 131K ~46/66 ~6 GiB 10.6 tok/s
Q4_K_XL at 80K, fitt 512 80K 64/66 1.10 GiB 51.5 tok/s
Practical Lessons
  • Never pin the GPU layer count (-ngl). Pinning disables auto-fit and caused a crash where the MTP draft context ran out of memory.
  • cache_reuse is auto-disabled by the server when using MTP with a single slot. The 8GB prompt cache and context checkpoints partially compensate.
  • There is a context cliff at about 100K tokens: accuracy degrades significantly around that mark even at full precision. That is why the context window is 80K and the compaction trigger sits at 85 percent of it (about 69.6K tokens).

Why the Q4_K_XL profile was chosen

Q4_K_XL was chosen because it matches the coding accuracy of higher-precision quantizations - both resolved 87.5 percent of a 16-task coding shootout - while decoding roughly five times faster (about 51.5 versus 10.6 tokens per second). Higher precision adds no measurable coding accuracy, so the faster profile is the daily driver.

The Agent Harness (DeepSeek Harness / dsh)

DeepSeek Harness (dsh) is a developer preview coding agent framework, like a self-hosted version of the cloud coding agents you may already know. It provides a web UI at http://127.0.0.1:3080 where you chat with the model, a preset system that controls which tools the agent can use, an MCP bridge that connects external tools (Unity, web fetcher, language server) to the model, hooks for pre-tool-use approval gates, automatic context compaction, and a plan mode where the agent proposes a plan before making changes.

Setup

  1. Install dsh globally (requires Node.js 24 or newer):
๐Ÿ–ฅ๏ธ Install dsh (PowerShell)
npm i -g @deepseek-ai/dsh
  1. Create a project profile by cloning the built-in web profile (the normal plugin add method fails because dsh ships with an unpublished dependency called dsh-fs-policy):
๐Ÿ–ฅ๏ธ Clone the web profile (PowerShell)
Remove-Item -Recurse -Force $env:USERPROFILE\.dsh\profiles\your-project-name
Copy-Item -Recurse $env:USERPROFILE\.dsh\profiles\web $env:USERPROFILE\.dsh\profiles\your-project-name
  1. Fix the profile name in package.json, since the clone inherits the name dsh-profile-web:
๐Ÿ–ฅ๏ธ Rename the profile (PowerShell)
$pkg = Get-Content "$env:USERPROFILE\.dsh\profiles\your-project-name\package.json" -Raw
$pkg = $pkg.Replace('dsh-profile-web', 'dsh-profile-your-project-name')
Set-Content "$env:USERPROFILE\.dsh\profiles\your-project-name\package.json" $pkg
  1. Configure the model provider in %USERPROFILE%\.dsh\settings.yaml:
๐Ÿ“„ %USERPROFILE%\.dsh\settings.yaml
llm-pi-ai:
  providers:
    local-qwen:
      apiKeyEnv: LOCAL_API_KEY
      api: openai-completions
      baseURL: http://127.0.0.1:8080/v1
      models:
        - id: qwen38-27b-local
          contextWindow: 81920
          maxTokens: 24576
agent-default-model:
  provider: local-qwen
  model: qwen38-27b-local
agent-presets:
  default: code-edit
  1. Set the API key as an environment variable. Any string works, since the local server does not validate it:
๐Ÿ–ฅ๏ธ API key + telemetry off (PowerShell)
[Environment]::SetEnvironmentVariable('LOCAL_API_KEY', 'local-no-auth', 'User')

# Disable telemetry (privacy requirement)
[Environment]::SetEnvironmentVariable('DSH_TELEMETRY_DISABLED', '1', 'User')
  1. Launch from the project root. The launch directory determines which AGENTS.md gets loaded:
๐Ÿ–ฅ๏ธ Launch dsh (PowerShell)
cd C:\Projects\your_project
dsh --profile your-project-name
Practical Lessons
  • The pnpm install path is broken for new profiles: dsh-fs-policy exists in the monorepo but was never published to npm. Cloning the working web profile sidesteps this entirely.
  • Presets are global, not per-project. Project-specific instructions belong in AGENTS.md at the project root, which dsh auto-injects into every session via its agent-instructions plugin.
  • Effort and sampler settings (temperature, top-p, reasoning effort) are not preset-encodable in the current dsh version (0.1.1-rc.2). Reasoning level is a per-model UI picker; samplers use the server's launch-flag defaults for all dsh traffic.
  • MCP tool subsets cannot be filtered per-preset. The MCP bridge exposes all tools from a connected server to all presets.
  • Hooks fail open: if the hook script cannot execute at all, the tool call proceeds silently. The only proof the approval gate works is seeing the approval prompt appear.

MCP Servers (External Tool Connections)

All MCP server entries go in %USERPROFILE%\.dsh\profiles\your-project-name\cordis.patch.yml, not in an mcp.json file (that file does not exist in Cordis-based profiles).

Unity MCP: scene and editor control

This connects the agent to the Unity Editor via the AI Connector plugin, running inside Unity on port 24152. It provides scene inspection, GameObject creation, modification and deletion, component access, material, shader and prefab operations, console log reading, dynamic C# execution, isolated screenshots, and the Unity test runner. Verified: asking it to explain the scene hierarchy returned the live scene tree.

๐Ÿ“„ cordis.patch.yml (Unity MCP)
- insert:
    - id: unity-mcp
      name: '@deepseek-ai/dsh-mcp-client'
      config:
        serverName: unity
        transport: streamable-http
        url: http://localhost:24152/mcp

Fetch MCP: web page retrieval

This lets the agent fetch web pages for documentation, using mcp-server-fetch, a Python tool run via uvx. Verified: fetching https://example.com returned the page content. The fetch server prints harmless warnings about Node.js not being found; it falls back to pure Python extraction.

๐Ÿ“„ cordis.patch.yml (Fetch MCP)
- insert:
    - id: fetch-mcp
      name: '@deepseek-ai/dsh-mcp-client'
      config:
        serverName: fetch
        transport: stdio
        command: uvx
        args: ['mcp-server-fetch']

Prerequisite: install uv, the Python package runner:

๐Ÿ–ฅ๏ธ Install uv (PowerShell)
powershell -ExecutionPolicy ByPass -c "irm https://astral.sh/uv/install.ps1 | iex"

Hooks: the approval gate before writes

This is the safety mechanism that prevents the agent from modifying your Unity scene without asking first.

Step 1: install the hooks bridge packages, pinned to your dsh version:

๐Ÿ–ฅ๏ธ Install hooks bridge (PowerShell)
cd $env:USERPROFILE\.dsh\profiles\your-project-name
pnpm add @deepseek-ai/dsh-hooks-claude-code@0.1.1-rc.2 @deepseek-ai/dsh-hook-protocol@0.1.1-rc.2

Step 2: create the approval script (ask-pretool.ps1):

๐Ÿ–ฅ๏ธ ask-pretool.ps1 (PowerShell)
'{"hookSpecificOutput":{"hookEventName":"PreToolUse","permissionDecision":"ask"}}' |
  Set-Content "$env:USERPROFILE\.dsh\profiles\your-project-name\ask-pretool.ps1"

Step 3: create the hooks config (hooks.json):

๐Ÿ“„ hooks.json
{
  "hooks": {
    "PreToolUse": [
      {
        "matcher": "mcp__unity__.*(create|delete|destroy|modify|set|add|remove|save|apply|import|move|rename|instantiate|edit|write|refresh|update|execute).*",
        "hooks": [
          {
            "type": "command",
            "command": "powershell -NoProfile -ExecutionPolicy Bypass -File \"%USERPROFILE%\\.dsh\\profiles\\your-project-name\\ask-pretool.ps1\""
          }
        ]
      }
    ]
  }
}

Step 4: register the hooks in cordis.patch.yml:

๐Ÿ“„ cordis.patch.yml (hooks)
- insert:
    - id: hooks-claude-code
      name: '@deepseek-ai/dsh-hooks-claude-code'
      config:
        configPath: '%USERPROFILE%\.dsh\profiles\your-project-name\hooks.json'
Verified Behavior

Creating a GameObject named "HITL-Probe" paused for approval. Deleting it also paused. Reading the console did not pause, which is correct: reads should stay ungated.

LSP: exact code navigation

The csharp-ls language server, built on Roslyn, lets the agent do compiler-grade find-references, go-to-definition, and go-to-implementation. It is read-only (no rename or refactor). The first query pays a 30 to 90 second solution load; subsequent queries are fast.

Step 1: install csharp-ls (requires .NET 10 SDK):

๐Ÿ–ฅ๏ธ Install csharp-ls (PowerShell)
dotnet tool install -g csharp-ls

Step 2: ensure Unity generates project files for embedded packages: Edit, Preferences, External Tools, tick "Embedded packages", then Regenerate project files. Verify a .slnx file exists at the project root.

Step 3: install the dsh LSP plugin packages:

๐Ÿ–ฅ๏ธ Install dsh LSP plugins (PowerShell)
cd $env:USERPROFILE\.dsh\profiles\your-project-name
pnpm add @deepseek-ai/dsh-lsp-stdio @deepseek-ai/dsh-lsp @deepseek-ai/dsh-tool-lsp

Step 4: register in cordis.patch.yml:

๐Ÿ“„ cordis.patch.yml (LSP)
- insert:
    - id: lsp
      name: '@deepseek-ai/dsh-lsp'
- insert:
    - id: lsp-stdio
      name: '@deepseek-ai/dsh-lsp-stdio'
      config:
        servers:
          csharp:
            command: '%USERPROFILE%\.dotnet\tools\csharp-ls.exe'
            args: ['-s', 'C:\Projects\your_project\YourProject.slnx']
            extensionToLanguage: { '.cs': csharp }
- insert:
    - id: tool-lsp
      name: '@deepseek-ai/dsh-tool-lsp'
Verified Behavior

goToImplementation on a base component class returned all implementing classes, correctly attributed to their source packages. Fast, with no errors. The LSP provides goToDefinition, findReferences, goToImplementation, and hover.

Presets (Agent Personas)

Three presets were created in %USERPROFILE%\.dsh\.agent-presets\, derived from dsh's shipped templates (standard and minimal) with cloud-dependent tools removed. Presets live in ~/.dsh/.agent-presets/<preset-id>/ with two files: agent.cordis.yml (tool composition) and preset.yml (display metadata). Presets are global, so they apply to all projects; project-specific rules go in AGENTS.md.

Preset Derived From Role
Code Edit standard The daily coding driver. Full toolset, plan mode enabled, compaction at 85 percent. The persona instructs the agent to follow the "Ponytail ladder": reuse existing code first, then platform built-ins, then the smallest new code.
Deep Debug standard Same toolset as Code Edit. The persona enforces "no hypothesizing without a red-capable feedback loop": the bug must be reproduced first before proposing fixes. Debug logging is tagged [DEBUG-xxxx] and removed after the fix.
Mechanical minimal Single-shot operations only: renames, moves, parameter changes. No thinking, no plan mode, no AGENTS.md injection. All local.

The agent-instructions plugin auto-injects AGENTS.md into sessions that include it (Code Edit and Deep Debug do; Mechanical does not).

What was removed from the standard template

Removed Tool Reason
tool-subagent-codex Cloud subagent provider (OpenAI), violates the offline requirement
tool-subagent-claude-code Cloud subagent provider (Anthropic), violates the offline requirement
tool-web Cloud web search, violates the offline requirement
delegation, tool-subagent, tool-subagent-control, tool-subagent-list-agents, tool-subagent-fork Subagents queue on the single-slot local server; the plan deprioritizes them
workflow-worker-thread, tool-workflow, tool-ralph Not in the plan; each tool's schema costs context tokens
Compaction Trigger Is Load-Bearing

Compaction is set to thresholdRatio: 0.85, meaning it triggers when the conversation reaches 85 percent of the model's context window (80K times 0.85, about 69.6K tokens). This is not a convenience: the model's accuracy degrades significantly past about 100K tokens.

Preset files to create

Each preset folder holds two files. The preset.yml display-metadata files are:

๐Ÿ“„ preset.yml ร— 3 (code-edit / deep-debug / mechanical)
# code-edit/preset.yml
name: Code Edit
description: Daily coding driver - full toolset, plan mode, compaction.
order: 10

# deep-debug/preset.yml
name: Deep Debug
description: Hard problems - root-cause debugging with a red-capable feedback loop.
order: 11

# mechanical/preset.yml
name: Mechanical
description: Single-shot operations - thinking off. Renames, moves, parameter changes.
order: 12

The agent.cordis.yml tool composition is derived from the shipped standard template (Code Edit and Deep Debug) or the minimal template (Mechanical), with the cloud/subagent rows removed as shown above. The only custom part is the persona text:

๐Ÿ“„ agent.cordis.yml - persona config (the custom part)
# code-edit (paste into persona.config.text of the derived `standard` preset)
You are a coding agent powered by {{model}}, working in {{cwd}}.
The workspace instructions (AGENTS.md chain) carry the project's stack, conventions, and
doctrine - read and follow them; more specific files take precedence.
Engineering discipline: reuse existing code first, then platform built-ins, then the smallest
new code that works. Fix root causes, never symptoms. Add one runnable check per non-trivial
logic. Never cut validation. If a change exceeds what you can verify, say so and stop.

# deep-debug (same template, different persona)
You are a debugging specialist powered by {{model}}, working in {{cwd}}.
The workspace instructions (AGENTS.md chain) carry the project's stack and conventions -
read and follow them. No hypothesizing without a red-capable feedback loop: reproduce the
failure first, by whatever means the project offers. Tag debug logging [DEBUG-xxxx] and
remove it after the fix. Fix root causes; never patch symptoms. If you cannot reproduce
the failure, say so and stop.

# mechanical (shipped `minimal` template; fixed prompt, no compaction, no AGENTS.md)
You are a mechanical-operations agent. Execute the requested operation exactly - renames,
moves, parameter changes, classifications. No redesign, no refactoring, no commentary
beyond the result. If the request is ambiguous, ask one precise question.

Project Instructions (AGENTS.md and CONTEXT.md)

The project's AGENTS.md at the Unity project root contains the operating rules the agent follows in every session:

  • Operating rules: smallest safe change first, preserve architecture, state uncertainty explicitly.
  • Project context: Unity 6 with three embedded packages (common, vehiclephysics, wheelcontroller).
  • The Ponytail ladder: reuse existing code, then Unity built-ins, then the vehicle physics and traffic APIs, then the smallest new code.
  • Validation rules: one runnable check per non-trivial logic, using the Unity Test Framework, edit-mode preferred.
  • Debugging doctrine: no hypothesizing without reproducing the bug first.
  • GitNexus: CLI-only in these sessions, run via node .gitnexus/run.cjs.
  • HITL: scene, prefab, and asset writes require approval.
  • Preset routing: Code Edit for daily work, Deep Debug for hard problems, Mechanical for single-shot tasks.
  • RAG sidecar instructions: how to call rag-query.ps1 for semantic code search.
  • LSP instructions: how to use the lsp tool for exact structural queries.
  • Skills pointer table: which SKILL.md file to read when a task matches.
๐Ÿ“„ AGENTS.md template
# AGENTS.md

## Operating rules

- Make the smallest safe change first.
- Preserve existing architecture unless a task requires refactor.
- Prefer explicit fixes over speculative rewrites.
- State uncertainty clearly.
- Keep code examples minimal and runnable.

## Repo workflow

- Read affected files before editing.
- When changing behavior, update or add tests when practical.
- Keep commits scoped to one concern.
- Do not rename files or modules without reason.

## Output preferences

- Be direct.
- Avoid filler.
- Use short explanations unless deeper detail is requested.
- Preserve exact commands, paths, error messages, and identifiers.

## Project Context

[Replace this section with your project's stack, packages, and conventions.]

## Agent Doctrine (dsh)

- Ponytail ladder, in order: reuse existing project code (search first) โ†’ Unity built-ins (Physics, NavMesh, WheelCollider, ScriptableObjects) โ†’ project/asset APIs โ†’ smallest new code that works.
- One runnable check per non-trivial logic; Unity Test Framework, edit-mode preferred. Never cut validation.
- Debugging: no hypothesizing without a red-capable feedback loop (batchmode test, console read, replayed input, or minimal scene harness). Tag debug logging [DEBUG-xxxx] and remove it after the fix.
- If a change exceeds what you can verify, stop and mark the ceiling with a `ponytail:` comment.
- GitNexus is CLI-only in these sessions: run it via PowerShell (`node .gitnexus/run.cjs ...`). If a needed query has no CLI equivalent, fall back to code search and say so.
- Scene, prefab, and asset writes require approval, so expect approval prompts and wait for them.
- Preset routing: Code Edit = daily work; Deep Debug = hard problems; Mechanical = single-shot ops.

## Exact code navigation (LSP)

An `lsp` tool is available (csharp-ls/Roslyn): goToDefinition, findReferences,
goToImplementation, hover. Use for exact structural queries: callers of X
(findReferences), class hierarchy (goToImplementation on the base class),
definitions, XML docs (hover). It needs a line:char position; grep for the
symbol first, then call lsp. Read-only; no rename/refactor ops.

## Semantic code search (RAG sidecar)

For concept-based code search ("where is X handled", "find code that does Y"),
run in PowerShell: C:\rag-sidecar\rag-query.ps1 "your query"
Returns ranked C# chunks with path:line citations. Use BEFORE writing new code
(check for existing implementations) and whenever grep keywords are unknown.
For exact structural queries (callers of X, class hierarchy) use the lsp tool (see above);
grep as fallback. After big code changes, re-index:
C:\rag-env\Scripts\python.exe C:\rag-sidecar\ingest_code.py C:\Projects\your_project

## Skills (read the SKILL.md file when the task matches)

| Task | Skill file |
|------|-----------|
| Sharpen a plan or design; produce ADRs + glossary | .dsh/skills/grill-with-docs/SKILL.md |
| Debugging: reproduce first, no hypothesizing | .dsh/skills/diagnosing-bugs/SKILL.md |
| New logic: red-green-refactor, Unity Test Framework | .dsh/skills/tdd/SKILL.md |
| Session ending / context nearly full | .dsh/skills/handoff/SKILL.md |

Supporting files (grilling, domain-modeling, tdd extras) sit in the same .dsh/skills tree.

## Session sign-off (CONTEXT.md refresh)

- At every sign-off, or at the end of a session, refresh CONTEXT.md, regenerating the compressed
  project briefing (architecture map, glossary, conventions, file:line anchors) so the next
  session starts current. Use rag-query.ps1 for concept discovery, the lsp tool and read for
  exact structure, and verify every fact against code. Facts only, no roadmap or aspirations.
- Do this ONLY if the user has explicitly approved it. Never regenerate CONTEXT.md
  automatically, on a timer, or by inference; an explicit user sign-off request is required first.

This file is auto-injected into every Code Edit and Deep Debug session by dsh's agent-instructions plugin.

CONTEXT.md: the project briefing file

A compressed briefing file at the project root gives future agent sessions the essential architecture, glossary, conventions, and file anchors without re-reading the entire codebase each time. It was generated autonomously by the agent itself, using the tools already available (RAG sidecar, LSP, grep, and file reads). This produces a verified glossary based on actual code rather than memory.

The generation prompt asks for a file of at most 400 lines covering the project architecture, a glossary of terms, conventions, and file:line anchors for the most-queried systems. Facts only, verified against code - no aspirations or roadmap. Paste this into a dsh session (Code Edit preset) after the RAG sidecar and LSP are live:

๐Ÿ’ฌ CONTEXT.md generation prompt
Generate CONTEXT.md at the project root - a compressed briefing for future agent
sessions, max 400 lines. Use rag-query.ps1 for concept discovery, the lsp tool for
exact structure, and read key files to verify. Include: (1) architecture map;
(2) glossary of project terms; (3) conventions - namespaces, asmdef boundaries,
logging, test layout; (4) file:line anchors for the most-queried systems.
Facts only, verified against code - no aspirations or roadmap.

A reference was added to AGENTS.md under Project Context: "Read CONTEXT.md first; keep it updated when architecture changes." Because it is auto-injected at the start of each session, it progressively shrinks the core feed, the stable context every session starts with, from roughly 55 to 60K tokens down to whatever CONTEXT.md covers.

The RAG Sidecar (Semantic Code Search)

The RAG sidecar lets the agent search code by meaning rather than by exact text. When the agent asks "where is engine braking handled?", the sidecar converts the question into a mathematical vector, finds the most similar code chunks, and returns them ranked by relevance.

The three components

  1. Qdrant, a vector database running in Docker. It stores every method, property, and class body as a separate chunk, parsed using tree-sitter, a code-aware parser.
  2. Harrier-oss-v1-0.6B, the embedding model. It converts text into 1024-dimensional vectors. It runs in-process via the Python library sentence-transformers on the CPU, inside the same Python venv as the retrieval scripts. There is no separate server and no HTTP port for the embedder. The model downloads automatically from the HuggingFace cache on first use. Documents are embedded without any prefix; queries carry an instruction prefix, as required by the model card.
  3. bge-reranker-v2-m3, a cross-encoder reranker. It also runs in-process via sentence-transformers. After Harrier retrieves 50 candidate chunks, the reranker re-reads each query and chunk pair jointly and re-scores them for a more precise final ranking.
No Separate Embedder Server

The final implementation uses the Python ML stack (torch CPU plus sentence-transformers) for both the embedder and the reranker. There is no second llama-server instance for embeddings and no port 8081. A rebuilder should not set one up.

Setup

Step 1: start Qdrant in Docker with a persistent volume:

๐Ÿ–ฅ๏ธ Start Qdrant (PowerShell)
docker run -d -p 6333:6333 -v C:\qdrant\storage:/qdrant/storage --name qdrant qdrant/qdrant
docker update --restart unless-stopped qdrant

Step 2: create a Python virtual environment and install dependencies:

๐Ÿ–ฅ๏ธ Python venv + dependencies (PowerShell)
python -m venv C:\rag-env
C:\rag-env\Scripts\python.exe -m pip install -r C:\rag-sidecar\requirements-rag.txt

Dependencies (requirements-rag.txt):

๐Ÿ“„ requirements-rag.txt
qdrant-client>=1.12
sentence-transformers>=3.1
tree-sitter>=0.23
tree-sitter-c-sharp>=0.23

This installs torch (CPU-only), sentence-transformers, qdrant-client, and tree-sitter with C# support. The first run of the ingestion or query script automatically downloads the Harrier model (about 1.2GB) and the reranker model (about 2.2GB) from HuggingFace. After that, everything runs fully offline.

Step 3: ingest your codebase:

๐Ÿ–ฅ๏ธ Ingest the codebase (PowerShell)
C:\rag-env\Scripts\python.exe C:\rag-sidecar\ingest_code.py C:\Projects\your_project

The ingestion script does the following:

  • Walks all .cs files in the project, respecting .cgcignore, which excludes Library/, Temp/, and Logs/.
  • Parses each file using the tree-sitter C# AST into method, property, and class level chunks.
  • Loads Harrier in-process and embeds each chunk (documents are embedded bare, with no instruction prefix).
  • Stores the vectors in Qdrant's code collection (1024 dimensions, cosine distance).
  • Tracks file modification times in .rag_index_state.json for incremental re-indexing: re-running only re-embeds files that changed, and purges deleted files.

The ignore file (.cgcignore), saved to the project root, controls what the ingestion walks:

๐Ÿ“„ .cgcignore
# C# code-search ignore file (read by ingest_code.py)
# Excludes generated, binary, and dependency folders from semantic indexing.
# Add your own patterns below.

node_modules/
venv/
.venv/
env/
.env/
dist/
build/
target/
out/
obj/
.git/
__pycache__/
*.png
*.jpg
*.jpeg
*.gif
*.svg
*.mp4
*.mp3
*.zip
*.tar
*.gz

Library/
Temp/
Logs/

The ingestion script itself (ingest_code.py), saved to C:\rag-sidecar\ingest_code.py:

๐Ÿ“„ C:\rag-sidecar\ingest_code.py
#!/usr/bin/env python3
"""Ingest a Unity project's C# sources into a Qdrant `code` collection.

Embedder: microsoft/harrier-oss-v1-0.6b (1024-dim dense, CPU).
Documents are embedded WITHOUT an instruction prefix (per model card);
queries carry the instruction at search time (see query.py).

Usage:
    python ingest_code.py C:\Projects\your_project
"""
import argparse
import fnmatch
import hashlib
import json
import sys
import time
from pathlib import Path

from qdrant_client import QdrantClient, models
from sentence_transformers import SentenceTransformer
from tree_sitter import Language, Parser
import tree_sitter_c_sharp as tscs

EMBED_MODEL = "microsoft/harrier-oss-v1-0.6b"
DENSE_DIM = 1024
STATE_NAME = ".rag_index_state.json"

MAX_CHUNK = 1000
WINDOW = 480
OVERLAP = 72

MEMBER_TYPES = {
    "method_declaration", "constructor_declaration", "property_declaration",
    "delegate_declaration", "operator_declaration", "indexer_declaration",
}
TYPE_TYPES = {
    "class_declaration", "struct_declaration", "interface_declaration",
    "enum_declaration", "record_declaration",
}
BUILTIN_IGNORES = {
    ".git", ".gitnexus", ".idea", ".vs", "Library", "Temp", "Logs",
    "obj", "bin", "Build", "Builds", "UserSettings", "MemoryCaptures",
}


def load_ignore_patterns(root: Path) -> list[str]:
    pats = []
    f = root / ".cgcignore"
    if f.exists():
        for line in f.read_text(encoding="utf-8", errors="replace").splitlines():
            line = line.strip()
            if line and not line.startswith("#"):
                pats.append(line)
    return pats


def is_ignored(rel: str, patterns: list[str]) -> bool:
    parts = rel.replace("\\", "/").split("/")
    if any(part in BUILTIN_IGNORES for part in parts[:-1]):
        return True
    rel_fwd = "/".join(parts)
    for pat in patterns:
        p = pat.rstrip("/")
        if fnmatch.fnmatch(rel_fwd, pat) or fnmatch.fnmatch(rel_fwd, pat + "/*"):
            return True
        if parts and (parts[0] == p or fnmatch.fnmatch(parts[0], p)):
            return True
    return False


def make_parser() -> Parser:
    return Parser(Language(tscs.language()))


def node_name(node) -> str:
    n = node.child_by_field_name("name")
    return n.text.decode("utf-8", "replace") if n else "<anon>"


def window_text(text: str) -> list[str]:
    if len(text) <= MAX_CHUNK:
        return [text]
    out, i = [], 0
    while i < len(text):
        out.append(text[i:i + WINDOW])
        i += WINDOW - OVERLAP
    return out


def class_context(node) -> str:
    cur = node.parent
    while cur is not None:
        if cur.type in TYPE_TYPES:
            return node_name(cur)
        cur = cur.parent
    return ""


def chunk_source(rel: str, source: bytes, parser: Parser) -> list[dict]:
    tree = parser.parse(source)
    chunks = []

    def emit(node, kind: str):
        body = source[node.start_byte:node.end_byte].decode("utf-8", "replace")
        cls = class_context(node)
        sym = f"{cls}.{node_name(node)}" if cls else node_name(node)
        start_line = node.start_point[0] + 1
        for k, piece in enumerate(window_text(body)):
            header = f"// {rel} :: {sym} ({kind})"
            chunks.append({
                "text": header + "\n" + piece,
                "path": rel,
                "symbol": sym,
                "kind": kind,
                "start_line": start_line,
                "end_line": node.end_point[0] + 1,
                "part": k,
            })

    def visit(node):
        if node.type in MEMBER_TYPES:
            emit(node, node.type.replace("_declaration", ""))
            return
        if node.type in TYPE_TYPES:
            text_len = node.end_byte - node.start_byte
            has_members = any(c.type in MEMBER_TYPES for c in node.children)
            if text_len <= MAX_CHUNK and not has_members:
                emit(node, node.type.replace("_declaration", ""))
                return
        for child in node.children:
            visit(child)

    visit(tree.root_node)
    if not chunks:
        body = source.decode("utf-8", "replace")
        for k, piece in enumerate(window_text(body)):
            chunks.append({
                "text": f"// {rel} :: <file>\n" + piece,
                "path": rel, "symbol": "<file>", "kind": "file",
                "start_line": 1, "end_line": body.count("\n") + 1, "part": k,
            })
    return chunks


def point_id(path: str, start_line: int, part: int) -> str:
    return hashlib.md5(f"{path}:{start_line}:{part}".encode()).hexdigest()


def main() -> int:
    ap = argparse.ArgumentParser()
    ap.add_argument("root")
    ap.add_argument("--qdrant", default="http://localhost:6333")
    ap.add_argument("--collection", default="code")
    args = ap.parse_args()

    root = Path(args.root).resolve()
    if not root.is_dir():
        print(f"not a directory: {root}", file=sys.stderr)
        return 1

    sys.stdout.reconfigure(encoding="utf-8")
    patterns = load_ignore_patterns(root)
    parser = make_parser()
    client = QdrantClient(url=args.qdrant, timeout=120)

    if not client.collection_exists(args.collection):
        client.create_collection(
            collection_name=args.collection,
            vectors_config=models.VectorParams(size=DENSE_DIM, distance=models.Distance.COSINE),
        )
        print(f"created collection '{args.collection}'")

    state_path = root / STATE_NAME
    state = json.loads(state_path.read_text()) if state_path.exists() else {}

    model = SentenceTransformer(EMBED_MODEL, device="cpu")

    seen, changed = set(), []
    for f in sorted(root.rglob("*.cs")):
        rel = f.relative_to(root).as_posix()
        if is_ignored(rel, patterns):
            continue
        seen.add(rel)
        mtime = f.stat().st_mtime_ns
        if state.get(rel) != mtime:
            changed.append((f, rel, mtime))

    stale = [p for p in state if p not in seen]
    for rel in stale:
        client.delete(
            collection_name=args.collection,
            points_selector=models.FilterSelector(filter=models.Filter(
                must=[models.FieldCondition(key="path", match=models.MatchValue(value=rel))])),
        )
        del state[rel]

    t0 = time.time()
    total_chunks = 0
    for f, rel, mtime in changed:
        chunks = chunk_source(rel, f.read_bytes(), parser)
        vecs = model.encode([c["text"] for c in chunks], batch_size=16,
                            normalize_embeddings=True, show_progress_bar=False)
        client.delete(
            collection_name=args.collection,
            points_selector=models.FilterSelector(filter=models.Filter(
                must=[models.FieldCondition(key="path", match=models.MatchValue(value=rel))])),
        )
        points = [
            models.PointStruct(
                id=point_id(rel, c["start_line"], c["part"]),
                vector=v.tolist(),
                payload=c,
            )
            for c, v in zip(chunks, vecs)
        ]
        client.upsert(collection_name=args.collection, points=points)
        state[rel] = mtime
        total_chunks += len(chunks)
        print(f"  {rel}: {len(chunks)} chunks", flush=True)

    state_path.write_text(json.dumps(state, indent=0))
    print(f"scanned {len(seen)} files | changed {len(changed)} | purged {len(stale)} "
          f"| upserted {total_chunks} chunks in {time.time() - t0:.1f}s")
    return 0


if __name__ == "__main__":
    raise SystemExit(main())

Step 4: create the query wrapper script (rag-query.ps1):

๐Ÿ–ฅ๏ธ rag-query.ps1 (PowerShell)
@'
param([Parameter(Mandatory=$true)][string]$Query, [int]$Top = 10)
& C:\rag-env\Scripts\python.exe C:\rag-sidecar\query.py "$Query" --top $Top
'@ | Set-Content C:\rag-sidecar\rag-query.ps1

The query.py script loads Harrier in-process, embeds the query with an instruction prefix, retrieves 50 candidate chunks from Qdrant using dense cosine search, then loads the reranker in-process and re-scores all 50 query and chunk pairs, returning the top N. A --no-rerank flag skips the reranker, which serves as an A/B testing arm for the evaluation harness.

The query script itself (query.py), saved to C:\rag-sidecar\query.py:

๐Ÿ“„ C:\rag-sidecar\query.py
#!/usr/bin/env python3
"""Query the Qdrant `code` collection.

Stage 1: harrier-oss-v1-0.6b dense search. Queries carry an instruction
prefix (required by the model card); documents were embedded without one.
Stage 2: bge-reranker-v2-m3 cross-encoder over the over-fetched candidates.

Usage:
    python query.py "how does the transmission calculate gear ratio"
"""
import argparse
import sys

from qdrant_client import QdrantClient
from sentence_transformers import SentenceTransformer

EMBED_MODEL = "microsoft/harrier-oss-v1-0.6b"
RERANK_MODEL = "BAAI/bge-reranker-v2-m3"
QUERY_INSTRUCTION = ("Instruct: Given a natural-language query, retrieve C# code "
                     "chunks that implement or explain the queried behavior\nQuery: ")
DISPLAY_CHARS = 600


def main() -> int:
    ap = argparse.ArgumentParser()
    ap.add_argument("query")
    ap.add_argument("--qdrant", default="http://localhost:6333")
    ap.add_argument("--collection", default="code")
    ap.add_argument("--fetch", type=int, default=50)
    ap.add_argument("--top", type=int, default=35)
    ap.add_argument("--no-rerank", action="store_true")
    args = ap.parse_args()

    sys.stdout.reconfigure(encoding="utf-8")
    client = QdrantClient(url=args.qdrant, timeout=120)
    if not client.collection_exists(args.collection):
        print(f"collection '{args.collection}' does not exist - run ingest_code.py first",
              file=sys.stderr)
        return 1

    model = SentenceTransformer(EMBED_MODEL, device="cpu")
    vec = model.encode([args.query], prompt=QUERY_INSTRUCTION,
                       normalize_embeddings=True)[0].tolist()

    res = client.query_points(
        collection_name=args.collection,
        query=vec,
        limit=args.fetch,
        with_payload=True,
    )
    points = res.points
    if not points:
        print("no results")
        return 0

    if args.no_rerank:
        ranked = [(p.score, p) for p in points[: args.top]]
    else:
        from sentence_transformers import CrossEncoder
        reranker = CrossEncoder(RERANK_MODEL)
        pairs = [(args.query, p.payload["text"]) for p in points]
        scores = reranker.predict(pairs)
        ranked = sorted(zip(scores, points), key=lambda t: t[0], reverse=True)[: args.top]

    for i, (score, p) in enumerate(ranked, 1):
        pl = p.payload
        body = pl["text"]
        if len(body) > DISPLAY_CHARS:
            body = body[:DISPLAY_CHARS] + "\n// ..."
        print(f"### {i}. {pl['path']}:L{pl['start_line']}-{pl['end_line']} "
              f"- {pl['symbol']} (score {float(score):.4f})\n")
        print(f"```csharp\n{body}\n```\n")
    return 0


if __name__ == "__main__":
    raise SystemExit(main())

Step 5: add a section to AGENTS.md so the agent knows the sidecar exists:

๐Ÿ“„ AGENTS.md (RAG sidecar section)
## Semantic code search (RAG sidecar)
For concept-based code search ("where is X handled", "find code that does Y"),
run in PowerShell: C:\rag-sidecar\rag-query.ps1 "your query"
Returns ranked C# chunks with path:line citations. Use BEFORE writing new code
(check for existing implementations) and whenever grep keywords are unknown.
For exact structural queries (callers of X, class hierarchy) use the lsp tool instead.
After big code changes, re-index:
  C:\rag-env\Scripts\python.exe C:\rag-sidecar\ingest_code.py C:\Projects\your_project

Measured performance

  • Ingestion: full project chunked and embedded on the CPU in about 282 seconds.
  • Query "how does the transmission calculate gear ratio": the top hit was the correct transmission gear-ratio method at score 0.9973, with the CVT gear ratio logic visible in the code.
  • Query "where is engine braking handled?": the agent autonomously used rag-query.ps1 following the AGENTS.md instructions, found the engine braking code, and produced a correct, cited answer in about 2 minutes at 43 tokens per second.
Practical Lessons
  • Harrier-oss-v1-0.6B is the correct embedder for this stack. It is a Microsoft open-source model (MIT license) with 1024-dimensional output, state of the art for its size class.
  • Harrier queries must carry an instruction prefix, per the model card, while documents are embedded bare. This asymmetry matters.
  • The .cgcignore file is read by ingest_code.py to exclude Library/, Temp/, and Logs/. Exclusions must be explicit there.
  • Both models run in-process through sentence-transformers, so no separate server is needed. This is simpler and more reliable than a separate embedder server approach.

The Skills Layer

Skills are instruction files (SKILL.md) that the agent reads on demand when a task matches the skill's purpose. They encode engineering workflows: plan before implementing, reproduce before debugging, test before refactoring, and summarize before handing off.

Six skills are installed from Matt Pocock's skills repository (github.com/mattpocock/skills) into the project's .dsh\skills folder:

๐Ÿ–ฅ๏ธ Install skills (PowerShell)
$dst = "C:\Projects\your_project\.dsh\skills"
$tmp = "$env:TEMP\pocock-skills.zip"
Invoke-WebRequest "https://github.com/mattpocock/skills/archive/refs/heads/main.zip" -OutFile $tmp
Expand-Archive $tmp "$env:TEMP\pocock-skills" -Force
$want = @("engineering\grill-with-docs","engineering\domain-modeling","engineering\diagnosing-bugs","engineering\tdd","productivity\grilling","productivity\handoff")
foreach ($s in $want) { Copy-Item -Recurse -Force "$env:TEMP\pocock-skills\skills-main\skills\$s" "$dst\$(Split-Path $s -Leaf)" }
Get-ChildItem $dst -Recurse -Directory -Filter agents | Remove-Item -Recurse -Force

This page's pinned dsh setup (0.1.1-rc.2) uses the AGENTS.md pointer table as its documented skill-loading mechanism (see below). Newer dsh source also documents an official filesystem skill provider with configurable roots, including a customSkillDirs option. Do not apply newer discovery behavior to the pinned setup without checking the installed dsh version.

Skill Purpose
grill-with-docs Sharpen a plan or design before implementing; produces ADRs and a glossary
domain-modeling Dependency of grill-with-docs; documents the domain model
grilling Dependency of grill-with-docs; the interactive questioning process
diagnosing-bugs Reproduce the bug first; no hypothesizing without a red-capable feedback loop
tdd Red-green-refactor workflow; Unity Test Framework, edit-mode preferred
handoff Session compaction and context transfer when a session ends or context is nearly full

The skills are registered in AGENTS.md through a pointer table that tells the agent which SKILL.md file to read when a task matches: grill-with-docs for sharpening a plan, diagnosing-bugs for debugging, tdd for new logic, and handoff when a session is ending or context is nearly full. Supporting files (grilling, domain-modeling, tdd extras) sit in the same skills tree.

How Skill Loading Works
  • Each skill is a SKILL.md file with YAML frontmatter containing name, description, and disable-model-invocation fields.
  • For the version used by this page (0.1.1-rc.2), the documented approach is the AGENTS.md pointer table: the agent reads the referenced SKILL.md when the task matches. Current dsh source documents a separate filesystem skill provider; its roots and configuration must be verified against the dsh version actually installed.
  • disable-model-invocation is preserved by construction: a skill only runs when the agent explicitly reads it on demand, not automatically on every session.
  • grill-with-docs is composite: its SKILL.md says to run a grilling session using the domain-modeling skill, so grilling and domain-modeling are required dependencies that must also be installed.

Unity-Generated Claude Skills and dsh

Unity's AI Connector can generate Claude-oriented files such as .claude/skills/gameobject-component-add/skill.md. Connecting Unity-MCP does not automatically make that file a dsh skill: MCP supplies tools, while a skill is an instruction file discovered by the agent runtime. The @deepseek-ai/dsh-hooks-claude-code package configured earlier in this page is an approval-hook bridge; it does not, by itself, load .claude/skills.

Do Not Assume the Unity File Is Discoverable

The example uses lowercase skill.md. The community compatibility plugin documented below explicitly describes .claude/skills/**/SKILL.md, so this page does not claim that lowercase filenames are accepted. Inspect the file's frontmatter and either preserve it only after a verified loader test, or copy/convert it to the exact filename and format required by the selected dsh provider. Keep the original Unity output unchanged until the conversion is confirmed.

There are two defensible integration paths. They must not be confused with the Zoo Code path above:

  1. Use the page's pinned baseline: copy or convert the useful instructions into .dsh/skills/<skill-name>/SKILL.md, retain the YAML frontmatter expected by this page, and add a row to the AGENTS.md pointer table. This is the least ambiguous route for the documented 0.1.1-rc.2 profile, but it is a migration rather than automatic reuse of .claude/skills.
  2. Use a current dsh filesystem provider or bridge: current official dsh source documents configurable skill roots, including customSkillDirs. The exact registration and supported filename must be checked against the installed release; this page does not invent a profile patch for an unpinned version.
Optional Community Compatibility Plugin

dsh-claude-compat is a third-party community plugin, not an official DeepSeek Harness baseline. Its README says it can expose project .claude/skills/**/SKILL.md through a dsh skill provider, and can separately read Claude commands, rules, and other integrations. It also states that rules are injected into the model context. Review and pin the source/version before installing it; a listing is not a security review, and trusted .claude/rules content can change agent behavior. Compatibility with this page's pinned 0.1.1-rc.2 setup is not established here, so this plugin is an optional experiment rather than a required step.

Verify the connection in layers: first confirm the skill file's name, filename, and frontmatter; then confirm that the selected dsh provider exposes the skill in its catalog or documented invocation surface; finally run a harmless request that should follow the skill. Separately test a read-only Unity-MCP call and a controlled write to confirm that the approval hook still behaves as configured. Skill visibility alone does not prove that Unity MCP is connected, and a successful Unity-MCP call does not prove that the skill was loaded.

The Windows Permission Fix

dsh's workspace-write sandbox can fail with a SetNamedSecurityInfoW failed (Win32 5) error on the project folder, meaning every shell command needs a full-access escalation. A common root cause: the project folder's owner SID belongs to a different Windows user (for example, a folder copied from another machine or an older Windows installation), so the current user account has no ownership and no write permission on the tree.

The fix, run from an elevated (administrator) PowerShell:

๐Ÿ–ฅ๏ธ ACL repair (elevated PowerShell)
takeown /F "C:\Projects\your_project" /R /D Y
icacls "C:\Projects\your_project" /reset /T /C /Q

takeown /R recursively takes ownership of the tree, and icacls /reset /T replaces every access rule with clean inherited permissions from the parent. After the fix, the sandbox grants write access cleanly. The first grantWrite call propagates the permission across the entire workspace tree, which can take a few minutes on a large Unity project but is a one-time cost.

The Complete Startup Sequence

To bring the entire stack up from a cold start:

  1. Qdrant auto-starts with Docker Desktop, since its restart policy is set to unless-stopped.
  2. The main model, in a separate window, left running:
๐Ÿ–ฅ๏ธ Start the model (PowerShell)
& "C:\llama-cpp\llama-server.exe" -m "C:\Qwen3.8-27B-UD-Q4_K_XL.gguf" --fit on -fitt 512 -c 81920 -fa on --jinja --cache-type-k q8_0 --cache-type-v q8_0 --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.7 --cache-reuse 256 -t 16 --parallel 1 --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0
  1. Unity Editor: open the project and ensure the AI Connector is running on port 24152.
  2. dsh, from the project root:
๐Ÿ–ฅ๏ธ Start dsh (PowerShell)
cd C:\Projects\your_project
dsh --profile your_custom_profile_name
No Embedder Server Needed

The Harrier embedder and the bge reranker do not need a separate server. They load in-process inside the Python venv when the ingestion or query script runs.

The verification probes

In a new dsh session with the Code Edit preset, run these checks:

  1. Ask any question. The model responds; watch the llama.cpp log for activity.
  2. "Fetch https://example.com". The fetch tool fires and returns the page content.
  3. "Read the Unity console". It runs without an approval prompt.
  4. "Create an empty GameObject named Verify-Probe". It pauses for approval; approve it, verify the object in the hierarchy, then delete it (it pauses again).
  5. "Where is engine braking handled?" The agent automatically uses rag-query.ps1 and returns a cited answer.
  6. "Use the lsp tool: goToImplementation on a base component class". It returns all implementing classes.

File Locations

Component Location
llama.cpp binary C:\llama-cpp\llama-server.exe
Main model (Qwen) C:\Qwen3.8-27B-UD-Q4_K_XL.gguf
dsh home %USERPROFILE%\.dsh\
dsh settings %USERPROFILE%\.dsh\settings.yaml
Profile config %USERPROFILE%\.dsh\profiles\your_custom_profile_name\cordis.patch.yml
Presets %USERPROFILE%\.dsh\.agent-presets\{code-edit,deep-debug,mechanical}\
HITL scripts %USERPROFILE%\.dsh\profiles\your_custom_profile_name\{ask-pretool.ps1,hooks.json}
RAG sidecar scripts C:\rag-sidecar\{ingest_code.py,query.py,rag-query.ps1,requirements-rag.txt}
RAG Python venv C:\rag-env\ (Python 3.12)
HF model cache (Harrier + reranker) %USERPROFILE%\.cache\huggingface\ (about 3.5GB)
csharp-ls %USERPROFILE%\.dotnet\tools\csharp-ls.exe
Skills C:\Projects\your_project\.dsh\skills\
AGENTS.md C:\Projects\your_project\AGENTS.md
CONTEXT.md C:\Projects\your_project\CONTEXT.md
Unity project C:\Projects\your_project
Qdrant volume C:\qdrant\storage

Extra Info

  • DFlash2: a faster speculative decoding method. Check llama.cpp pull request #27342; if it is merged, it can replace MTP for more speed.
  • CONTEXT.md maintenance: regenerate or update it when the architecture changes significantly.

7b Local OCR-to-Markdown Studio (MinerU + Gradio - bonus)

This bonus setup turns your machine into a local browser-based PDF OCR application that converts scanned PDFs into structured Markdown using MinerU and a small Gradio wrapper. After the one-time setup, you run one command (or double-click a launcher), drop PDFs into a browser page, and receive Markdown, JSON, and extracted images locally.

No source PDF is uploaded by this wrapper. The Gradio server binds only to 127.0.0.1, so it is reachable from the same machine only. It complements the RAG sidecar above: converted Markdown becomes the corpus for a future docs Qdrant collection.

Requirements

Target platform: Windows 10 or Windows 11 x64. The Gradio wrapper concept also works on Linux and macOS, but GPU installation differs by platform and vendor - do not apply the Windows CUDA instructions to Linux, macOS, or AMD GPUs.

Hardware Supported route Recommendation
NVIDIA GPU CUDA-enabled PyTorch, optional CUDA Toolkit for the local hybrid backend Best for throughput and complex document understanding
AMD GPU AMD/ROCm route (normally Linux) Follow MinerU's AMD/ROCm docs; do not install CUDA
Apple Silicon MinerU's MPS route Follow MinerU's macOS guidance
No compatible GPU CPU pipeline backend Works, but slower and less capable on difficult layouts

Plan for at least 20 GB free SSD space (models, packages, caches, output) and 16 GB RAM (32 GB preferable). Use the current MinerU documentation as the source of truth for VRAM, because requirements change between releases.

Python version

On Windows use Python 3.10, 3.11, or 3.12 x64. Do not use 3.13 or newer unless MinerU's current documentation explicitly says it is supported.

Terminal - Windows PowerShell
py --list
py -3.12 --version

Project layout

Create a dedicated folder in a user-writable location. Avoid protected locations such as C:\Program Files, because the project needs to write logs and output files.

Example layout
OCR-Markdown-Studio\
  run.py
  requirements.txt
  install.ps1
  run.cmd
  parse-with-gpu.bat
  .venv\
  logs\
  results\
๐Ÿ–ฅ๏ธ Create the project folder (PowerShell)
New-Item -ItemType Directory -Force "$HOME\Apps\OCR-Markdown-Studio" | Out-Null
Set-Location "$HOME\Apps\OCR-Markdown-Studio"

Project files

Create requirements.txt. Do not put a generic torch dependency here - GPU-enabled PyTorch is hardware-specific, and a generic install may replace a working accelerator build with a CPU-only build.

๐Ÿ“„ requirements.txt
gradio>=5,<6

Create install.ps1:

๐Ÿ“„ install.ps1
$ErrorActionPreference = "Stop"

if (!(Get-Command py -ErrorAction SilentlyContinue)) {
    throw "Python Launcher was not found. Install Python 3.10, 3.11, or 3.12 x64 first."
}

py -3.12 -m venv .venv

& .\.venv\Scripts\Activate.ps1

python -m pip install --upgrade pip
pip install -r requirements.txt

Write-Host ""
Write-Host "Base UI installation complete." -ForegroundColor Green
Write-Host "Next: install the accelerator-specific PyTorch build, then MinerU."

If you use Python 3.10 or 3.11 instead, replace both occurrences of 3.12 with the installed compatible version.

Create run.cmd:

๐Ÿ“„ run.cmd
@echo off
setlocal

if not exist ".venv\Scripts\python.exe" (
  echo Missing virtual environment.
  echo Run install.ps1 first.
  pause
  exit /b 1
)

".venv\Scripts\python.exe" run.py

After installation, a user can double-click run.cmd to open the local browser application.

The Gradio app (run.py)

Create run.py with this complete content:

๐Ÿ“„ run.py
from __future__ import annotations

import os
import shutil
import subprocess
import sys
import time
from pathlib import Path

import gradio as gr


APP_DIR = Path(__file__).resolve().parent
DEFAULT_OUTPUT = APP_DIR / "results"
LOG_DIR = APP_DIR / "logs"

DEFAULT_OUTPUT.mkdir(exist_ok=True)
LOG_DIR.mkdir(exist_ok=True)

# Require previously downloaded local MinerU models during normal conversion.
os.environ["MINERU_MODEL_SOURCE"] = "local"


def find_mineru() -> str:
    executable = shutil.which("mineru") or shutil.which("mineru.exe")

    if executable:
        return executable

    candidate = Path(sys.executable).parent / "mineru.exe"
    if candidate.exists():
        return str(candidate)

    raise RuntimeError(
        "MinerU is not installed in this virtual environment. "
        "Activate .venv and install MinerU first."
    )


def find_pdfs(uploaded_files: list[str] | None, folder: str | None) -> list[Path]:
    pdfs: list[Path] = []

    for file_path in uploaded_files or []:
        path = Path(file_path)
        if path.exists() and path.suffix.lower() == ".pdf":
            pdfs.append(path)

    if folder:
        folder_path = Path(folder).expanduser()
        if folder_path.is_dir():
            pdfs.extend(folder_path.rglob("*.pdf"))

    return list(dict.fromkeys(pdfs))


def read_log(log_path: Path, limit: int = 12000) -> str:
    if not log_path.exists():
        return ""

    return log_path.read_text(
        encoding="utf-8",
        errors="replace",
    )[-limit:]


def convert(files, folder, output_folder, quality, progress=gr.Progress()):
    documents = find_pdfs(files, folder)

    if not documents:
        yield "Select PDF files or provide a folder containing PDFs.", ""
        return

    try:
        mineru = find_mineru()
    except RuntimeError as error:
        yield str(error), ""
        return

    output_path = Path(output_folder or DEFAULT_OUTPUT).expanduser().resolve()
    output_path.mkdir(parents=True, exist_ok=True)

    log_path = LOG_DIR / f"conversion-{time.strftime('%Y%m%d-%H%M%S')}.log"
    summary: list[str] = []

    for index, pdf in enumerate(documents, start=1):
        progress(
            (index - 1) / len(documents),
            desc=f"Starting {pdf.name}",
        )

        # Omitting -b lets MinerU select its configured/default local backend.
        command = [
            mineru,
            "-p", str(pdf),
            "-o", str(output_path),
        ]

        if quality == "high":
            command.extend(["--effort", "high"])

        with log_path.open("a", encoding="utf-8") as log:
            log.write("\n\n$ " + subprocess.list2cmdline(command) + "\n")

            process = subprocess.run(
                command,
                stdout=log,
                stderr=subprocess.STDOUT,
                text=True,
            )

        status = (
            f"Complete: {pdf.name}"
            if process.returncode == 0
            else f"Failed (exit {process.returncode}): {pdf.name}"
        )

        summary.append(status)

        progress(
            index / len(documents),
            desc=status,
        )

        yield "\n".join(summary), read_log(log_path)

    yield (
        "\n".join(summary)
        + f"\n\nResults: {output_path}"
    ), read_log(log_path)


def open_output_folder(output_folder: str) -> str:
    folder = Path(output_folder or DEFAULT_OUTPUT).expanduser().resolve()
    folder.mkdir(parents=True, exist_ok=True)

    if sys.platform.startswith("win"):
        os.startfile(folder)
    elif sys.platform == "darwin":
        subprocess.Popen(["open", str(folder)])
    else:
        subprocess.Popen(["xdg-open", str(folder)])

    return f"Opened: {folder}"


with gr.Blocks(
    title="OCR Markdown Studio",
    theme=gr.themes.Soft(),
) as app:
    gr.Markdown(
        "# OCR Markdown Studio\n"
        "Local PDF OCR and Markdown conversion using MinerU."
    )

    with gr.Row():
        files = gr.File(
            label="Drop PDF files",
            file_count="multiple",
            file_types=[".pdf"],
            type="filepath",
        )

        folder = gr.Textbox(
            label="Or convert every PDF in a folder",
            placeholder=r"C:\Documents\Scans",
        )

    with gr.Row():
        output_folder = gr.Textbox(
            label="Output folder",
            value=str(DEFAULT_OUTPUT),
        )

        quality = gr.Radio(
            label="Quality",
            choices=[
                ("Normal", "medium"),
                ("High, includes image/chart analysis", "high"),
            ],
            value="medium",
        )

    with gr.Row():
        convert_button = gr.Button(
            "Convert to Markdown",
            variant="primary",
        )

        open_button = gr.Button("Open output folder")

    status = gr.Textbox(
        label="Status",
        lines=6,
    )

    logs = gr.Code(
        label="Conversion log",
        language="shell",
        lines=18,
    )

    convert_button.click(
        convert,
        inputs=[files, folder, output_folder, quality],
        outputs=[status, logs],
    )

    open_button.click(
        open_output_folder,
        inputs=[output_folder],
        outputs=[status],
    )


if __name__ == "__main__":
    app.launch(
        server_name="127.0.0.1",
        server_port=7860,
        inbrowser=True,
        show_error=True,
    )
The critical privacy line

server_name="127.0.0.1" binds the UI to the local loopback interface only. Do not change it to 0.0.0.0 unless you deliberately want other devices on the network to reach it.

First-time installation

Open a normal PowerShell window in the project directory. Avoid PowerShell ISE for interactive download tools, because progress prompts can behave poorly there.

๐Ÿ–ฅ๏ธ First install (PowerShell)
Set-ExecutionPolicy -Scope Process Bypass
.\install.ps1
.\.venv\Scripts\Activate.ps1

The prompt should begin with (.venv).

Install PyTorch correctly

Install the correct PyTorch build as required by MinerU's current documentation, then verify it. A generic dependency resolver may install a CPU-only Torch build even on a GPU system. Run this check after each major install step:

๐Ÿ–ฅ๏ธ Torch accelerator check (PowerShell)
python -c "import torch; print('Torch:', torch.__version__); print('CUDA:', torch.version.cuda); print('Available:', torch.cuda.is_available()); print('GPU:', torch.cuda.get_device_name(0) if torch.cuda.is_available() else 'No CUDA GPU')"
  • NVIDIA: use the PyTorch selector at pytorch.org and pick the Windows/Pip/CUDA command for your driver and CUDA build.
  • AMD: CUDA commands are wrong - follow MinerU's AMD/ROCm instructions.
  • CPU-only: install standard CPU PyTorch and use the pipeline backend rather than a local CUDA VLM backend.
NVIDIA example pattern

This is an example pattern, not a universal command:

pip uninstall -y torch torchvision torchaudio

pip install torch torchvision torchaudio `
  --index-url https://download.pytorch.org/whl/cuXXX

Replace cuXXX with the CUDA wheel channel the current PyTorch selector recommends, then re-run the check and proceed only if it prints True.

Install MinerU

With the virtual environment still active:

๐Ÿ–ฅ๏ธ Install MinerU (PowerShell)
python -m pip install --upgrade pip
pip install -U "mineru[all]"
mineru --help
mineru-models-download --help

Re-check GPU after MinerU - essential for NVIDIA users. If a previously working CUDA build becomes CPU-only, reinstall the matching CUDA-enabled Torch/TorchVision wheels and re-run the check; do not reinstall everything blindly.

NVIDIA CUDA Toolkit note

A CUDA-enabled PyTorch wheel often supplies enough runtime libraries for Torch. But some local MinerU high-accuracy configurations use a backend such as LMDeploy/TurboMind that also requires the full NVIDIA CUDA Toolkit and the Windows CUDA_PATH variable.

If a local hybrid/VLM run fails with Can not find $env:CUDA_PATH, install the CUDA Toolkit version compatible with your PyTorch/MinerU stack from NVIDIA's official CUDA Toolkit Archive, then verify in a fresh PowerShell:

๐Ÿ–ฅ๏ธ Verify CUDA Toolkit (PowerShell)
$env:CUDA_PATH
where.exe nvcc
nvcc --version

Do not install the CUDA Toolkit on AMD systems merely because an error message mentions CUDA.

Download models once

Download the document-processing models before locking to local-only mode:

๐Ÿ–ฅ๏ธ Download all model groups (PowerShell)
mineru-models-download --source auto --model_type all

Sources: auto, huggingface, modelscope. Model types: pipeline, vlm, all. If one group fails, retry only that group from the alternate source, for example mineru-models-download --source modelscope --model_type vlm. To start with standard OCR/layout conversion only: mineru-models-download --source auto --model_type pipeline.

Test MinerU directly

Always validate MinerU at the command line before diagnosing a Gradio-wrapper issue.

๐Ÿ–ฅ๏ธ Direct CLI test (PowerShell)
New-Item -ItemType Directory -Force "C:\Temp\mineru-output" | Out-Null
mineru -p "C:\Path\To\document.pdf" -o "C:\Temp\mineru-output"
Backend Practical meaning When to use it
Default / local hybrid engine High-quality local processing if the accelerator engine is installed After validating GPU/backend configuration
pipeline General, broader-compatibility parsing First fallback for CPU use or accelerator issues
VLM/hybrid HTTP client Connects to a separately running compatible model server Advanced deployment only

For a scanned PDF and a standard OCR-first test, explicitly request OCR where supported (mineru -p "..." -o "..." -b pipeline -m ocr). For a proven local hybrid backend, a higher-effort mode may be available (--effort high) - use it only after a normal test works.

Force local-only operation

Only after all required models download and a direct conversion succeeds, configure MinerU to use local model files:

๐Ÿ–ฅ๏ธ Set MINERU_MODEL_SOURCE (PowerShell)
[Environment]::SetEnvironmentVariable("MINERU_MODEL_SOURCE", "local", "User")

Close all PowerShell windows and open a new one, then verify with [Environment]::GetEnvironmentVariable("MINERU_MODEL_SOURCE", "User") - expected output: local. This tells MinerU to use previously downloaded local models; the wrapper itself has no cloud endpoint or API-key configuration.

Run the browser application

๐Ÿ–ฅ๏ธ Launch the app (PowerShell)
.\run.cmd
# or: .\.venv\Scripts\python.exe .\run.py

The app opens at http://127.0.0.1:7860. Normal use: drag PDFs into the file box (or enter a folder path), pick an output folder, start with Normal quality, click Convert to Markdown, then inspect the generated Markdown before relying on it.

Typical output layout
results\
  DocumentName\
    DocumentName.md
    DocumentName_content_list.json
    images\

Sequential GPU scheduling

OCR uses the GPU and must never run while the local LLM server is loaded - both compete for VRAM. The launcher below parses a PDF or folder, then restarts the LLM automatically, so the whole detour is one double-click. Save it as parse-with-gpu.bat in the project folder:

๐Ÿ“„ parse-with-gpu.bat
@echo off
setlocal
REM ============================================================
REM  GPU PDF parse (MinerU) - sequential GPU scheduling helper
REM  Run ONLY while the local LLM server is stopped.
REM  Parses the given PDF/folder, then restarts the LLM server.
REM  Usage: parse-with-gpu.bat "C:\path\to\file.pdf-or-folder"
REM ============================================================

set INPUT=%~1
if "%INPUT%"=="" set /p INPUT=PDF file or folder to parse:
if "%INPUT%"=="" (
  echo Nothing given. Aborting.
  pause
  exit /b 1
)

REM Refuse to run if the LLM server is still holding the GPU
curl -s -o nul -w "%%{http_code}" http://127.0.0.1:8080/health 2>nul | findstr /b "200" >nul
if not errorlevel 1 (
  echo.
  echo  WARNING: llama-server is still responding on port 8080.
  echo  Close the llama-server window first, then re-run this file.
  echo.
  pause
  exit /b 1
)

echo Parsing: %INPUT%
"%USERPROFILE%\Apps\OCR-Markdown-Studio\.venv\Scripts\mineru.exe" -p "%INPUT%" -o "%USERPROFILE%\Apps\OCR-Markdown-Studio\results"
set RC=%ERRORLEVEL%

echo.
if %RC%==0 (echo Parse complete.) else (echo Parse FAILED with exit code %RC%.)

echo Restarting LLM server (FAST profile)...
start "llama-server FAST" cmd /k "C:\llama-cpp\llama-server.exe" -m "C:\Qwen3.8-27B-UD-Q4_K_XL.gguf" --fit on -fitt 512 -c 81920 -fa on --jinja --cache-type-k q8_0 --cache-type-v q8_0 --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.7 --cache-reuse 256 -t 16 --parallel 1 --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0

echo.
echo Results: %USERPROFILE%\Apps\OCR-Markdown-Studio\results
pause
exit /b %RC%

The bat guards the dangerous case: it health-checks port 8080 first and refuses to parse while the model is loaded, so even a confused sequence cannot crash on VRAM contention.

AGENTS.md: PDF parsing section

Add this next to the RAG sidecar section of the project AGENTS.md so the agent knows the staged stop-LLM โ†’ parse โ†’ auto-restart loop:

๐Ÿ“„ AGENTS.md - PDF parsing section
## PDF parsing (MinerU, GPU - sequential GPU scheduling)

PDF-to-Markdown OCR runs on the GPU and must NEVER run while the local LLM server
is loaded (VRAM contention). You cannot run it yourself - stopping the LLM stops
you. When PDF parsing is needed:

1. Tell the user: "Stop the LLM server (close the llama-server window), then run
   this file - it parses and restarts the LLM automatically:"
   %USERPROFILE%\Apps\OCR-Markdown-Studio\parse-with-gpu.bat ""
2. Say exactly what you want parsed and why, so the user can verify the output.
3. Wait - do not continue until the user confirms the server is back up.
4. Output lands in %USERPROFILE%\Apps\OCR-Markdown-Studio\results\ as Markdown.

Troubleshooting

  • No suitable Python runtime found: install a supported x64 Python (3.10-3.12) and update install.ps1 to request it (for example py -3.11 -m venv .venv).
  • torch.cuda.is_available() is False: run nvidia-smi, then inspect the Torch build. If it contains +cpu or reports CUDA: None, install the CUDA-enabled PyTorch build for the NVIDIA path.
  • MinerU changed CUDA Torch to CPU Torch: note the installed versions with pip show mineru lmdeploy torch torchvision torchaudio, reinstall a compatible CUDA-enabled Torch/TorchVision set, then re-verify GPU detection.
  • Can not find $env:CUDA_PATH: a backend expects the full NVIDIA CUDA Toolkit. Install a compatible Toolkit, open a new terminal, and confirm CUDA_PATH and nvcc exist.
  • Model download fails: retry the failed group with the alternate source. Do not set MINERU_MODEL_SOURCE=local until the necessary model groups are present.
  • Gradio opens but conversion fails: run the exact PDF through the CLI first. If CLI works but Gradio fails, check the app's Conversion log and the logs directory, and confirm run.py is started with the project .venv, not a system-wide Python.
  • Slow conversion: first run loads models into RAM/VRAM; CPU parsing is slower than GPU; use Normal before high-effort; try a short PDF first; monitor GPU with nvidia-smi -l 2.

Updating safely

Do not casually run pip install -U mineru torch torchvision in a working environment - a resolver may replace accelerator-specific builds. Safer: back up the project and outputs, record pip freeze > requirements-working.txt, update one component at a time, re-run the Torch accelerator check, test one representative PDF with the CLI, then use the Gradio app.

Security and privacy checklist

  • Keep the server binding as 127.0.0.1; do not expose port 7860 through a firewall rule or port forwarding.
  • Do not use cloud/VLM HTTP client modes unless you deliberately accept external data transmission.
  • Download models first, then enable MINERU_MODEL_SOURCE=local.
  • Store outputs on an encrypted disk if the PDFs contain sensitive data.
  • Review generated Markdown - OCR may misread low-quality scans, handwriting, stamps, tables, numbers, or unusual layouts. Use human review for documents where errors matter.

Sources

Keyboard Shortcuts

Search
CtrlK
Search (alt)
/
Close search
Esc
Next result
โ†“
Prev result
โ†‘
Jump to result
Enter
Top of page
Home