antibot-print

command module
v0.0.0-...-f9ec972 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Jun 2, 2026 License: MIT Imports: 21 Imported by: 0

README

antibot

Print the antibot vendors protecting a site by matching its HTTP response against a single regex.

Install

macOS, Linux, WSL:

curl -fsSL https://raw.githubusercontent.com/albinstman/antibot-print/main/install.sh | bash

Windows (PowerShell):

irm https://raw.githubusercontent.com/albinstman/antibot-print/main/install.ps1 | iex

macOS: binaries are unsigned — clear the quarantine flag once with xattr -d com.apple.quarantine ~/.local/bin/antibot.

Usage

Print vendors:

$ antibot https://example.com
cloudflare

Or pipe response from curl:

$ curl -isS https://example.com | antibot
cloudflare

Add -c to report only vendors actively serving a challenge or block, not mere presence:

$ antibot -c https://example.com
cloudflare

Add -n to fetch with Go's default fingerprint. Surfaces vendors that challenge bots but pass real browsers:

$ antibot -n https://example.com
perimeterx

Use -p to impersonate a different browser (e.g. chrome_146, firefox_135). The default is chrome_148; chrome_147 and chrome_148 reuse chrome_146's TLS/HTTP-2 fingerprint (unchanged since 146) with only the User-Agent bumped:

$ antibot -p firefox_135 https://example.com
cloudflare

Add -d for diagnostics:

$ antibot -d https://example.com
request:
  url:    https://example.com
  mode:   browser (profile chrome_148)
response:
  status: 200
  bytes:  4521
  redirects:
    https://example.com (301) -> https://www.example.com/ (200)
detection (presence):
  cloudflare
    ← H:set-cookie:__cf_bm=
detection (challenge):
  (none)

Use -D to add the normalized view and full raw response:

$ antibot -D https://example.com > debug.txt

Use -r to print only the raw fetched response (status line, headers, body — like curl -i -L), with no detection output. The exit code still reflects detection (0 if a vendor was found, 1 if none), so it doubles as a fingerprinted curl:

$ antibot -r https://example.com > response.txt

To run the regex yourself instead of the binary, see Language integration.

Tip: the two fingerprints answer different questions. The default browser fetch shows what a real browser gets served. -n shows what an obvious bot gets served. Some vendors only reveal themselves to one or the other. When in doubt, run both.

Update to the latest release:

$ antibot update
antibot: updated abc1234 → def5678

Agent skill

SKILL.md is an agent skill that teaches a coding agent (e.g. Claude Code) to use antibot. To install it, paste install-skill.md to your agent.

Language integration

Go
package main

import (
	"fmt"
	"net/http"
	"net/http/httputil"
	"os"
	"regexp"
	"strings"
)

func normalize(raw []byte) string {
	t := strings.ReplaceAll(strings.ReplaceAll(string(raw), "\r\n", "\n"), "\r", "\n")
	head, body, _ := strings.Cut(t, "\n\n")
	out := []string{}
	for i, line := range strings.Split(head, "\n") {
		if i == 0 {
			if m := regexp.MustCompile(`HTTP/[\d.]+\s+(\d{3})`).FindStringSubmatch(line); m != nil {
				out = append(out, "S:"+m[1])
			}
			continue
		}
		if k, v, ok := strings.Cut(line, ":"); ok {
			out = append(out, "H:"+strings.ToLower(strings.TrimSpace(k))+":"+strings.TrimSpace(v))
		}
	}
	if len(body) > 8*1024*1024 {
		body = body[:8*1024*1024]
	}
	out = append(out, "B:"+strings.NewReplacer("\n", " ", "\t", " ").Replace(body))
	return strings.Join(out, "\n")
}

func main() {
	resp, err := http.Get("https://example.com")
	if err != nil {
		panic(err)
	}
	defer resp.Body.Close()
	raw, _ := httputil.DumpResponse(resp, true) // raw status line + headers + body

	regexText, _ := os.ReadFile("antibot.re2.txt")
	re := regexp.MustCompile(strings.TrimSpace(string(regexText)))

	norm := normalize(raw)
	names := re.SubexpNames()
	for _, m := range re.FindAllStringSubmatch(norm, -1) {
		for i, g := range m {
			if i > 0 && g != "" && names[i] != "" {
				fmt.Println(names[i])
			}
		}
	}
}
Python
import re, urllib.request, urllib.error  # or: import re2 as re

def normalize(raw, cap=8 * 1024 * 1024):
    text = raw.decode("latin-1").replace("\r\n", "\n").replace("\r", "\n")
    head, _, body = text.partition("\n\n")
    lines = head.split("\n")
    out = []
    m = re.match(r"HTTP/[\d.]+\s+(\d{3})", lines[0])
    if m:
        out.append("S:" + m.group(1))
    for line in lines[1:]:
        name, sep, val = line.partition(":")
        if sep:
            out.append(f"H:{name.strip().lower()}:{val.strip()}")
    out.append("B:" + body[:cap].replace("\n", " ").replace("\t", " "))
    return "\n".join(out)

def main(url):
    req = urllib.request.Request(url, headers={"User-Agent": "antibot"})
    try:
        r = urllib.request.urlopen(req)
    except urllib.error.HTTPError as e:
        r = e  # a 403/429 challenge is still the response we want
    raw = f"HTTP/1.1 {r.status} {r.reason}\r\n{r.headers}\r\n".encode("latin-1") + r.read()

    rx = re.compile(open("antibot.re2.txt").read().strip())
    norm = normalize(raw)
    for v in sorted({g for m in rx.finditer(norm) for g, val in m.groupdict().items() if val}):
        print(v)

if __name__ == "__main__":
    main("https://example.com")
JavaScript
import fs from "node:fs";

function normalize(raw, cap = 8 * 1024 * 1024) {
  const text = Buffer.from(raw).toString("latin1").replace(/\r\n|\r/g, "\n");
  const i = text.indexOf("\n\n");
  const [head, body] = i < 0 ? [text, ""] : [text.slice(0, i), text.slice(i + 2)];
  const lines = head.split("\n");
  const out = [];
  const m = lines[0].match(/^HTTP\/[\d.]+\s+(\d{3})/);
  if (m) out.push("S:" + m[1]);
  for (const line of lines.slice(1)) {
    const c = line.indexOf(":");
    if (c >= 0) out.push("H:" + line.slice(0, c).trim().toLowerCase() + ":" + line.slice(c + 1).trim());
  }
  out.push("B:" + body.slice(0, cap).replace(/[\r\n\t]/g, " "));
  return out.join("\n");
}

const res = await fetch("https://example.com");
const headers = [...res.headers].map(([k, v]) => `${k}: ${v}`).join("\r\n");
const raw = Buffer.from(`HTTP/1.1 ${res.status} ${res.statusText}\r\n${headers}\r\n\r\n` + (await res.text()));

// JS regex differs from RE2: translate named groups, drop (?m), scope (?i:).
const pat = fs.readFileSync("antibot.re2.txt", "utf8").trim()
  .replace(/^\(\?m\)/, "").replace(/\(\?P</g, "(?<").replace(/\(\?i:/g, "(?:");
const rx = new RegExp(pat, "gmi");

const norm = normalize(raw);
const vendors = new Set();
for (const m of norm.matchAll(rx))
  for (const [name, v] of Object.entries(m.groups || {})) if (v !== undefined) vendors.add(name);
console.log([...vendors].sort());

These normalizers cover the common single-response case; Normalize in main.go is the reference (redirect chains, byte handling).

Project structure

signatures/<vendor>.json       source of truth: {vendor, signals:[RE2], challenge?:[RE2 subset]}
main.go                        normalize, compile, detect; embeds the regexes
fetch.go                       direct-fetch path: browser-fingerprinted HTTP via tls-client
debug.go                       `--debug` diagnostic report (request, detection, raw response)
update.go                      `antibot update` + throttled "update available" notifier
main_test.go                   smoke tests + artifact-sync guard
fetch_test.go                  fetch helpers: redirect/Location/header-order tests
update_test.go                 update helpers: asset name / checksum parsing tests
antibot.re2.txt                compiled presence regex (embedded, usable standalone)
antibot-challenge.re2.txt      compiled challenge-only regex (embedded, usable standalone)
install.sh                     curl | bash installer (downloads a release binary)
.github/workflows/release.yml  build 5 platforms on push to main -> rolling "latest" release
vendor/                        vendored Go deps

Build from source

go build -o antibot .   # embeds antibot.re2.txt + antibot-challenge.re2.txt
go test ./...           # smoke tests + artifact-sync check
go run . compile        # regenerate both .re2.txt artifacts from signatures/

To add or change a vendor, edit a signatures/<vendor>.json (each signal prefixed S:/H:/B:, valid RE2, vendor-specific) and run go run . compile. Pushing to main recompiles the regexes and rebuilds every platform binary into the rolling latest release, so the signature files are the only thing you maintain.

License

MIT

Documentation

Overview

--debug path: instead of just naming vendors, dump a diagnostic of one response — how it was fetched, the redirect chain, every vendor matched in the active tier with the exact text that triggered it, and (with -D) the normalized view the regex runs against plus the full raw response. It prints to stdout in place of the slug list.

Direct-fetch path: retrieve a URL ourselves and feed the response to detection.

Two fingerprints are available, because the request shape changes what antibots reveal in opposite directions:

  • browser (default): a Chrome TLS/HTTP-2 fingerprint and header set, so evasion-aware WAFs (Akamai, DataDome, …) reveal their cookies/scripts.
  • naive (--naive): Go's stdlib client with its default Go TLS/HTTP fingerprint and no browser headers, so challenge-on-suspicion vendors (e.g. PerimeterX) that pass clean browsers silently still serve their block page to an obvious bot.

Either way we follow redirects manually, capturing each hop's raw response and concatenating them (like `curl -i -L`) so intermediate Set-Cookie/headers — where challenges are frequently planted — survive into Normalize's multi-block parsing.

Command antibot names the antibot/WAF/CAPTCHA vendor(s) protecting a site from a static HTTP response read on stdin. It is a single self-contained binary: the compiled regex is embedded, and Go's stdlib regexp IS RE2 (linear-time, no backreferences/lookaround), so nothing else is bundled or required at runtime.

Self-update: an explicit `antibot update` command plus a throttled, TTY-only notifier. The notifier never changes the binary — it prints a one-line stderr hint at most once a day and only when stderr is a terminal, so piped/scripted/CI use stays silent and offline. `antibot update` is the only thing that replaces the binary, and it verifies the download against the release's SHA256SUMS first.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL