parallel-web-extract
Extract content from multiple URLs in parallel, token-efficiently.
- Handles webpages, articles, PDFs, and JavaScript-heavy sites with a single command
- Runs in a forked context to minimize token overhead compared to built-in WebFetch
- Supports batch extraction of multiple URLs with optional focus objectives
- Requires
parallel-cliinstallation and authentication; outputs extracted content as markdown to a local file for follow-up queries
URL Extraction
Extract content from: $ARGUMENTS
Command
Choose a short, descriptive filename based on the URL or content (e.g., vespa-docs, react-hooks-api). Use lowercase with hyphens, no spaces. Substitute it into the command inline — $FILENAME is a placeholder, not a shell variable.
Pass each requested URL as a separate quoted positional argument, up to 20 per call. Do not collapse multiple URLs into one quoted $ARGUMENTS string or use eval to split them. Construct arguments directly from the requested URLs. For example:
parallel-cli extract "https://docs.parallel.ai/integrations/cli" "https://docs.parallel.ai/integrations/cursor-marketplace" --json -o "/tmp/parallel-docs.json"
-o saves JSON. Use a .json extension and inspect an existing path before use because Extract overwrites it. Read the saved file as authoritative; stdout may truncate and human-readable output previews only part of the content. Do not treat a stale file as a successful response after a failed call.
Options if needed: