KnockOutEZ/wigolo

▲ 31 stars today★ 5,344⑂ 432

The go-to web for your AI coding agent — local-first search, fetch, crawl & research over MCP. No API keys, no cloud, $0/query. Public beta.

About KnockOutEZ/wigolo

KnockOutEZ/wigolo is an open-source project on GitHub, mainly written in TypeScript. The go-to web for your AI coding agent — local-first search, fetch, crawl & research over MCP. No API keys, no cloud, $0/query. Public beta. It currently holds 5,344 stars and 432 forks with 0 open issues, and was last pushed on an unknown date (repository created unknown).

Project Overview

Git Homed tracks it on the Today's Trending board, currently at rank #78 with 31 new stars today.

GitHub Repository Details

Repository KnockOutEZ/wigolo · default branch - · size 0 KB · watchers 0 · source: GitHub REST API and repository README

README

https://github.com/KnockOutEZ/wigolo/blob/HEAD/wigolo — the go-to web for your agent

Local-first web intelligence for AI agents — no keys, no cloud, no metered bill.

works with  Claude Code · Cursor · Codex · Gemini CLI · OpenCode · VS Code · Windsurf · Zed · Antigravity
and beyond  LangChain · CrewAI · LlamaIndex · Vercel AI SDK · n8n & self-hosted agents · any MCP client · plain REST

npm npm downloads GitHub stars CI node MCP license status follow on X Discord

https://github.com/KnockOutEZ/wigolo/blob/HEAD/wigolo on Trendshift https://github.com/KnockOutEZ/wigolo/blob/HEAD/KnockOutEZ%2Fwigolo | Trendshift

Quickstart · Tools · Why wigolo · Discord · Sponsors · Benchmark · Docs · Examples · Feedback · FAQ

Join the community on Discord — questions, help, and what's being built next.

New features and updates ship steadily. Follow @yourtowhid on X for all of it and new ways to use wigolo, and reach out there for collaborations or feedback · also on LinkedIn

---

wigolo gives an AI agent one surface for everything web-related: search, fetch, crawl, extract, cache, find-similar, research, and autonomous gather loops. It runs wherever your agent runs — as an MCP server next to your coding agent, as a REST/MCP endpoint on the box where your self-hosted agents live, or embedded through an SDK inside your own app. The core tools need no API keys, nothing it touches leaves ~/.wigolo/, and no bill grows with how much your agent thinks.

https://github.com/KnockOutEZ/wigolo/blob/HEAD/wigolo demo — Claude Code answering a live web question through wigolo, no API keys

Quickstart

npx wigolo init                              # set up the local engine — any system
npx wigolo init --agents=claude-code,cursor  # …or set up + wire your day-to-day agents in one command

Requires Node ≥ 20 and ~1.5 GB of free disk on macOS, Linux, or Windows. Bare init sets up the local engine: it downloads the browser engine and on-device models, runs a health check, and reports each component. Adding --agents wires the named agents in the same run, so a coding agent you use daily is ready in one command.

init is unattended by default, so it's safe in scripts and CI, and any setup problem surfaces right here in the per-component report, before your agent's first call. Search, fetch, crawl, extract, cache, and find-similar work with no API key. Check it's healthy anytime:

npx wigolo doctor

To remove everything cleanly, run npx wigolo config --uninstall --yes. You can also paste the installation guide into any AI assistant and let it do the setup; it's written to be self-contained.

Recommended — a free key for research & agent

Search, fetch, crawl, extract, cache, and find-similar are fully keyless. research, agent, and search format=answer use an LLM to write the synthesized, cited answer. Without one they hand back a raw brief and evidence for your agent to assemble. A free Gemini key turns that into a finished answer:

export WIGOLO_LLM_PROVIDER=gemini
export GEMINI_API_KEY=      # grab one at aistudio.google.com/apikey — the free tier is plenty

Any provider works (anthropic · openai · groq), or stay fully local and keyless with WIGOLO_LLM_PROVIDER=ollama (or any OpenAI-compatible URL). Set it in your shell or your agent's MCP env block. Providers, models, and the keyless local-model ladder are in the configuration guide.

What your agent gets back

Every search result is evidence the agent can act on. It carries a verbatim excerpt pinned to its exact position in the source, a citation ID the agent can quote, and a score it can inspect (abridged real shape):

{
  "results": [{
    "title": "Logical replication - PostgreSQL docs",
    "url": "https://www.postgresql.org/docs/current/logical-replication.html",
    "excerpt": "Logical replication is a method of replicating data objects…",
    "citation_id": "src-1",
    "source_span": { "start": 1042, "end": 1305 },          // byte-exact provenance
    "evidence_score": { "final": 0.86, "semantic": 0.91, "lexical": 0.78, "engine_consensus": 3 }
  }],
  "citations": [{ "id": "src-1", "url": "…" }],
  "freshness_signal": { "published": "2026-05-12", "confidence": "high" }
}

Weak results get flagged as junk by wigolo's own scorer. Failed engines are reported and stale cache is labeled, so the agent always knows what it's standing on. Full response contracts per tool are in the tools reference.

Tools

| Tool | What it does | |------|--------------| | 🔎 search | Multi-engine web search (18 direct adapters) with rank fusion, ML reranking, and an explainable per-result score. Pass a query array for parallel breadth. Scope by domain and time range, match an exact phrase, or return image results. | | 📄 fetch | Load one URL through a tiered router that auto-escalates from plain HTTP to a headless browser engine on anti-bot challenges or SPA shells. Clean markdown + metadata + links. Handles PDFs, a single-heading section, authenticated sessions, and page actions (click / type / scroll / screenshot). | | 🕸️ crawl | Multi-page crawl — BFS, DFS, sitemap, or map-only. Per-domain rate limits, robots.txt respect, boilerplate dedup. | | 🧩 extract | Structured data from a page: tables, metadata, JSON-LD, brand identity, named schemas (Article / Recipe / Product / …), or any custom JSON Schema. | | 💾 cache | Query everything already seen — keyword or hybrid semantic. Plus stats, clear, and change detection. | | 🧲 find_similar | Pages similar to a URL or a concept, via 3-way fusion of keyword + semantic + live web. | | 🧠 research | Decompose a question → fan out sub-queries → fetch sources → synthesize a cited report (or a structured brief the host LLM writes from). | | 🤖 agent | Autonomous gather loop: plan → search → fetch → extract → synthesize, with a step log, time budget, and optional output schema. | | 🔁 diff + ⏱️ watch | See exactly what changed on a page since last visit; re-check on demand and deliver changes to a webhook. |

Every tool also runs from the terminal (wigolo search "…" --json), from an interactive shell with NDJSON piping (wigolo shell), over REST, and through the SDKs — CLI reference. Per-tool guides with the full parameter set are in docs/tools.md; runnable examples are in examples/.

Why it's different

wigolo isn't a free stand-in for the paid tools — it's built to match them. It's a focused web layer for your agents: an MCP and REST surface they call directly, with the search and extraction quality the paid services charge for. What separates it:

Here's what one real result looks like, dissected. It includes the failed engine and the weak result, because those are part of the answer too:

https://github.com/KnockOutEZ/wigolo/blob/HEAD/Anatomy of a wigolo result: explainable score decomposition, live engine telemetry, surfaced degradation, self-flagged junk — one real query, captured live

Sponsors

Thank you to the sponsors below, who help keep wigolo maintained and free for everyone to use. Their support goes straight into the work.

https://github.com/KnockOutEZ/wigolo/blob/HEAD/TestMu AI

TestMu AI (formerly LambdaTest) is the world's first full-stack agentic AI quality engineering platform, trusted by 18,000+ enterprises.

wigolo is free for all and is meant to stay that way. If you or your company would like to help keep it maintained, there's room for more sponsors — reach out at [email protected], or see SPONSORS.md for the terms. A one-off via Buy Me a Coffee is welcome too.

Benchmark

All four tools converged on the same core answer, and only one of them handed back verbatim, byte-pinned evidence while doing it.

One cold query ran live inside a single Claude Fable 5 session, fanned out to four web tools on equal footing (built-in WebSearch, wigolo, Tavily, Exa), and was judged by the agent on the evidence alone. All four converged on the same answer and the same top source, so the parity is demonstrated on-screen. wigolo alone returned verbatim excerpts pinned to byte-offset source spans, an explainable score decomposition, and live per-engine telemetry, and its own scorer flagged two weak results as junk. The cloud tools earn their place too: Exa rendered the official docs' comparison matrix in full. Run your own query and you'll see the same shape.

https://github.com/KnockOutEZ/wigolo/blob/HEAD/wigolo vs built-in WebSearch, Tavily, and Exa on one real query, driven by Claude Fable 5

How it compares

| | wigolo | Firecrawl | Exa | Tavily | |---|:---:|:---:|:---:|:---:| | Multi-engine web search | ✅ | ✅ | ✅ | ✅ | | Fetch & structured extraction | ✅ | ✅ | ✅ | ✅ | | Whole-site crawl & map | ✅ | ✅ | — | ✅ | | Verbatim excerpts pinned to byte-offset source spans | ✅ | — | — | — | | Explainable per-result score decomposition | ✅ | — | — | — | | Persistent local memory — re-query instantly, offline | ✅ | — | — | — | | Query data stays on your machine | ✅ | — | — | — | | API key / account | none | required | required | required | | Cost per query | $0 | metered | metered | metered |

Feature standing as of July 2026 — check each vendor's docs for current state.

That last row compounds, because agents ask in bursts:

https://github.com/KnockOutEZ/wigolo/blob/HEAD/The meter: a metered cloud API's cost climbs with every query while wigolo stays flat at zero dollars — illustrative pricing

Beyond your editor

The same ten tools serve every kind of agent, over whichever surface fits: MCP for coding agents, REST for everything else, SDKs to embed, and framework wrappers to drop in.

REST API — wigolo serve

One process exposes a plain-JSON REST API next to the MCP transport. No MCP client needed, just curl:

wigolo serve                          # 127.0.0.1:3333 — loopback is open; off-loopback requires a token

curl -sX POST http://127.0.0.1:3333/v1/search \ -H 'Content-Type: application/json' \ -d '{"query":"local-first software","max_results":5}'

POST /v1/{tool} covers all ten tools, GET /openapi.json is the OpenAPI 3.1 contract, and /mcp + /sse serve remote MCP clients from the same port. Bind past loopback and a bearer token is required, so the server fails closed by default. Point n8n, a Hermes-style assistant, or any self-hosted agent at it. → REST API

SDKs — TypeScript & Python

Thin, typed clients with an embedded local mode that finds or starts the daemon for you. No separate serve step.

TypeScriptnpm install wigolo-sdk (zero-dep; Node / Bun / Deno / edge):

import { createLocalClient } from 'wigolo-sdk/local';

const { client, close } = await createLocalClient(); // reuse a running daemon, or spawn one const res = await client.search({ query: 'local-first web search', max_results: 5 }); console.log(res.results.map((r) => r.title)); await close(); // stops the daemon only if this call spawned it

Pythonpip install wigolo (standard library only; sync + async):

from wigolo import local_client

with local_client() as client: # reuse a healthy daemon, or spawn one res = client.search(query="local-first web search", max_results=5) for r in res["results"]: print(r["title"], r["url"])

SDKs & embedded mode

Framework integrations

Drop wigolo's tools into the framework you already use. You get the full ten-tool surface, including the cache / find_similar / research / agent that most framework web-tools don't ship:

| Framework | Package | What you get | |-----------|---------|--------------| | LangChain | wigolo-langchain | each tool as a BaseTool, plus a BaseRetriever over search / find_similar for RAG | | CrewAI | wigolo-crewai | wigolo_tools() → hand the set to any crew | | LlamaIndex | wigolo-llamaindex | a BaseReader that loads fetched / crawled / searched pages as documents | | Vercel AI SDK | wigolo-vercel-ai-sdk | tool factories for generateText / streamText, edge-friendly |

Framework integrations

Docker

# stdio MCP — wire it into any MCP client as command: docker
docker run -i --rm -v wigolo-data:/data ghcr.io/knockoutez/wigolo

HTTP server for remote / multi-client use

docker run -p 3333:3333 -v wigolo-data:/data \ -e WIGOLO_API_TOKEN=a-long-random-secret \ ghcr.io/knockoutez/wigolo serve --host 0.0.0.0

The slim image lazy-loads models into the volume; :full preinstalls the browser engine. Also on Docker Hub as towhid69420/wigolo. → installation & all channels

Agent skills

An 11-pack skill catalog teaches your coding agent to drive each tool well. It's installed by init and managed with wigolo skills add|list|remove. → skills

One note for self-hosters: some challenge-protected sites score IP reputation, so a datacenter IP won't clear walls a home connection would. wigolo labels those failures, and the self-hosting guide covers the opt-in proxy answer.

Star history

https://github.com/KnockOutEZ/wigolo/blob/HEAD/wigolo GitHub stars over time

Refreshed daily from the GitHub API. Add a ⭐ if wigolo is useful to you.

Architecture

A single Node process speaks MCP (JSON-RPC over stdio). Everything heavy is local and lazy-loaded, so a zero-key install pays nothing for the parts it isn't using.

flowchart TD
    A["🤖 AI agent
any MCP client · REST · SDK"] A -->|MCP over stdio| B["wigolo
10 tools · dynamic instructions
in-process browser pool + cache + models"]

B --> C{"Tool layer"} C --> T1["search · fetch · crawl · extract"] C --> T2["cache · find_similar · research · agent"]

T1 --> F["⚙️ Fetch router
tiered escalation, learned per domain"] T1 --> S["⚙️ Search
18 engines → rank fusion → ML rerank
explainable evidence score"] T2 --> DB[("🗄️ Local cache
keyword + vector index")] T2 --> ML["🧠 On-device ML
embeddings + reranker"]

F -.->|optional| LLM["☁️ LLM
synthesis only · opt-in"] S -.->|optional| SX["🔀 Aggregator backend
opt-in legacy / hybrid"]

F --> WEB["🌍 Public web"] S --> WEB

style B fill:#7c3aed,stroke:#5b21b6,color:#fff style WEB fill:#0ea5e9,stroke:#0369a1,color:#fff style DB fill:#1e293b,stroke:#334155,color:#fff style LLM stroke-dasharray: 5 5 style SX stroke-dasharray: 5 5

Configuration

A clean install works out of the box. Three settings raise output quality:

# 1. Synthesis — the biggest lever (research / agent / search-answer write real prose)
export WIGOLO_LLM_PROVIDER=gemini                   # or anthropic / openai / groq / ollama (keyless)
export GEMINI_API_KEY=

2. Wider retrieval funnel

export WIGOLO_SEARCH=hybrid # core engines + aggregator fallback export WIGOLO_GITHUB_TOKEN=... # GitHub code search 10 → 30 req/min

3. Land more fetches, stay warm

export WIGOLO_TLS_TIER=auto # per-domain learned fetch hardening export WIGOLO_EAGER_WARMUP=1 # pay the ~1s model load up front

Per-call habits that pay off: query arrays (["a","b","c"]) for parallel breadth · search_depth: "deep" for queries that matter · include_domains as a hard filter for docs lookups. The full reference covers every environment variable, config-file key, search backend, cache TTL, and serve limit; it's in the configuration guide.

Docs & examples

docs/ — the complete manual: getting started · installation & channels · configuration · tools reference · CLI & shell · REST API · SDKs & integrations · self-hosting · agent skills · plugins · troubleshooting & FAQ · privacy & security

examples/ — runnable, each with a README (and most with a terminal recording): one-shot CLI, NDJSON shell pipelines, REST via curl, TypeScript & Python SDKs, Vercel AI SDK tools, pointing self-hosted n8n at a remote wigolo, watch-with-webhook, and writing your own search-engine plugin. The docs are also rendered on the site at knockoutez.github.io/wigolo/docs.

Beta & feedback

wigolo is in public beta. Everything documented here works and is held to a 7,600-test suite; it's stable, and beta is about the polish bar. It stays beta until enough people have used it, kicked it, and starred it that calling it v1 means something. Your feedback shapes what comes next, and every report is read, usually the same day:

If wigolo earns a place in your setup, three things keep it going: a ⭐ star (it's how open source gets found), a ☕ coffee (there's no paid tier and never will be), or an email that goes straight to the one developer who wrote the code.

Troubleshooting

wigolo doctor names any broken component and the exact env var or command that fixes it; wigolo doctor --fix repairs the common cases, and wigolo verify health-checks every component. A component failing during init doesn't break wigolo: init still exits 0, and core search / fetch / crawl / extract / cache work with no models and n

GitHub Stars & Activity

5,344Stars
432Forks
0Open issues
TypeScriptLanguage

GitHub Popularity

GitHub stars5,344
Forks432
Open issues0
Primary languageTypeScript
License-
Stars gained today31
Created-
Last pushed-

Trending History

Daily boardrank #78 · ▲ 31 stars

Related GitHub Projects

1

anthropics / claude-code

TypeScript★ 146,951⑂ 23,997▲ 415 stars
2

OpenCut-app / OpenCut

TypeScript★ 89,948⑂ 8,888▲ 126 stars
3

ZuodaoTech / everyone-can-use-english

TypeScript★ 38,103⑂ 5,221▲ 436 stars
4

supermemoryai / supermemory

TypeScript★ 30,670⑂ 2,676▲ 392 stars
5

browserbase / stagehand

TypeScript★ 24,631⑂ 1,698▲ 146 stars
6

vercel-labs / json-render

TypeScript★ 17,059⑂ 907▲ 585 stars
7

ahmedkhaleel2004 / gitdiagram

TypeScript★ 16,720⑂ 1,275▲ 285 stars
8

Open-Dev-Society / OpenStock

TypeScript★ 16,537⑂ 2,122▲ 752 stars

More Trending Repositories