Getting started

research web

Nine operations cover URL discovery, page retrieval, crawling, structured extraction, saved indexes, history and site management.

Sign in with orc login and select an accessible --workspace. API keys and agent delegation are unsupported. Unset ORCHESTOR_API_KEY when using user login because it takes precedence over stored credentials. Existing beta access requirements apply.

orc login
orc research web --help
orc research web map --url https://example.com --fallback none --workspace YOUR_WORKSPACE_ID --format json

--fallback none disables provider fallback. Map defaults to firecrawl.

web crawl

orc research web crawl [flags]

POST /v1/research/crawl

FlagRequiredType and constraints
--fallback—string; enum=["none", "firecrawl"]; default="none"
--formats—array; default=["markdown"]; minItems=1
--limit—integer; default=5; minimum=1; maximum=10
--max-age—integer; default=86400; minimum=0; maximum=86400
--max-depth—integer; default=1; minimum=0; maximum=3
--url✓string; maxLength=2048

web extract

orc research web extract [flags]

POST /v1/research/extract

FlagRequiredType and constraints
--fallback—string; enum=["none", "firecrawl"]; default="none"
--max-age—integer; default=86400; minimum=0; maximum=86400
--mode—string; enum=["selectors", "jsonld"]; default="selectors"
--selectors—array; default=[]; maxItems=20
--urls✓array; minItems=1; maxItems=10

web history get

orc research web history get [flags]

GET /v1/research/history

FlagRequiredType and constraints
--url✓string; maxLength=2048

web index get

orc research web index get [flags]

GET /v1/research/index

FlagRequiredType and constraints
--url✓string; maxLength=2048

web map

orc research web map [flags]

POST /v1/research/map

FlagRequiredType and constraints
--fallback—string; enum=["none", "firecrawl"]; default="firecrawl"
--limit—integer; default=100000; minimum=1; maximum=100000
--max-age—integer; default=86400; minimum=0; maximum=86400
--sitemap—string; enum=["include", "only", "skip"]; default="include"
--url✓string; maxLength=2048

web scrape

orc research web scrape [flags]

POST /v1/research/scrape

FlagRequiredType and constraints
--fallback—string; enum=["none", "firecrawl"]; default="none"
--formats—array; default=["markdown"]; maxItems=2
--max-age—integer; default=86400; minimum=0; maximum=86400
--urls✓array; minItems=1; maxItems=10

web sites list

orc research web sites list [flags]

GET /v1/research/sites

No command-specific inputs.

web sites create

orc research web sites create [flags]

POST /v1/research/sites

FlagRequiredType and constraints
--domain✓string; maxLength=253

web sites discover

orc research web sites discover [flags]

POST /v1/research/sites/discover

FlagRequiredType and constraints
--domain✓string; maxLength=253

Inputs and acquisition bounds

--urls and --formats accept comma-separated values. --max-age uses seconds and --max-depth controls link depth. URLs cannot contain credentials or arbitrary ports.

orc research web scrape --urls https://example.com,https://example.com/docs --formats markdown --max-age 0 --workspace YOUR_WORKSPACE_ID
orc research web crawl --url https://example.com --limit 5 --max-depth 1 --fallback none --workspace YOUR_WORKSPACE_ID
orc research web extract --urls https://example.com --mode jsonld --workspace YOUR_WORKSPACE_ID
orc research web extract --urls https://example.com --mode selectors --selectors '[{"name":"title","selector":"h1","multiple":false}]' --workspace YOUR_WORKSPACE_ID
orc research web index get --url https://example.com --workspace YOUR_WORKSPACE_ID
orc research web history get --url https://example.com --workspace YOUR_WORKSPACE_ID
orc research web sites create --domain example.com --workspace YOUR_WORKSPACE_ID
orc research web sites discover --domain example.com --workspace YOUR_WORKSPACE_ID
orc research web sites list --workspace YOUR_WORKSPACE_ID
  • Map discovers up to 100,000 URLs from sitemaps, llms.txt and links. It does not retrieve every page body.
  • Scrape and extract accept up to 10 URLs per call. Crawl traverses up to 10 pages and depth 3. CLI crawl is link traversal; the site screen's Crawl batches known URLs up to 1,000 pages.
  • Extract parses CSS selectors or JSON-LD without LLM inference. Selectors mode requires 1–20 uniquely named selectors. attribute reads an attribute; multiple: true returns multiple matches. Use --stdin for complex JSON inputs.
  • To persist scraped bodies while returning metadata only, pass "formats": [] through JSON input.
  • Sites create only registers a domain; it starts no acquisition or schedule. Discover returns up to 30 unregistered sites linked from a homepage without registering them, saving captures or invoking paid fallback.

Responses, persistence and errors

JSON retains the full API response inside data. For acquisition, inspect each URL's status in data.data, plus data.truncated, data.warnings and data.usage. HTTP 200 may include partial failures or truncation. Usage counts requests, not currency. Sites and history return data.data; index returns data.result and data.fetchedAt, null when no saved index exists.

Acquisition persists content. Shareable anonymous captures may be reused for up to 24 hours. Query-bearing URLs, non-shareable responses and Firecrawl captures remain workspace-scoped. History returns up to 20 versions readable for 30 days. --max-age 0 requests fresh acquisition. Caller authentication cookies are never sent to source sites.

Map runs synchronously for up to 90 seconds; other acquisition operations are bounded to 40 seconds. Robots denials and private network restrictions are retained. Map's default Firecrawl fallback can invoke the external provider when native sitemap discovery is incomplete. Other primitives only allow provider fallback on blocked responses when explicitly selected.

--dry-run previews a request without sending or saving it. --raw removes the CLI envelope. Automatic traversal with --page-all is unavailable. Exit codes: success 0, API/authentication/network failure 1, argument error 2. For 401/403, check user login, Workspace and access requirements. Inspect per-URL outcomes before retrying partial failures.