Menu

AI agents

Give LLMs and agents eyes on the web with screenshots, Markdown and vision, and point them at these docs.

Pages as Markdown

format=markdown renders a page in a real browser and returns its content as Markdown: JavaScript-rendered content included, scripts and styles removed. It's a compact way to put a web page into an LLM's context.

curl -X POST "https://shotkit.net/api/take" \
  -H "X-Access-Key: YOUR_ACCESS_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url":"https://example.com","format":"markdown","block_cookie_banners":true,"selector":"main"}' \
  --fail-with-body -o shot.md

selector narrows it to the part that matters and saves tokens. For pages behind a login, pass cookies or an Authorization header.

Screenshots for vision models

Vision models read screenshots well. Keep them small: a 1280-wide viewport capture as JPG is plenty, and image_width scales it down further:

curl -X POST "https://shotkit.net/api/take" \
  -H "X-Access-Key: YOUR_ACCESS_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url":"https://example.com","format":"jpg","image_quality":70,"image_width":1024,"block_cookie_banners":true,"block_chats":true}' \
  --fail-with-body -o shot.jpg

To get the picture and the text in one render, add metadata_content with response_type=json; see HTML and Markdown.

Ask a question about the page

With your OpenAI key and a vision_prompt, the screenshot is sent to an OpenAI vision model and its answer comes back with the result:

curl -X POST "https://shotkit.net/api/take" \
  -H "X-Access-Key: YOUR_ACCESS_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url":"https://example.com","response_type":"json","openai_api_key":"sk-YOUR_OPENAI_KEY","vision_prompt":"Is there a cookie banner on this page? Answer yes or no.","vision_max_tokens":10}' \
  --fail-with-body -o shot.json
{
  "screenshot_url": "https://shotkit.net/cdn/files/…/….jpg",
  "vision": { "completion": "No" }
}

Your key is used for this request only and never stored. OpenAI errors fail the request with vision_error.

Give your agent these docs

Every page of this documentation is plain Markdown for agents:

In the Markdown version, examples are cURL commands, so an agent can run them directly. The Copy page button at the top of each page copies the same Markdown, ready to paste into a chat.

A system prompt line that works well:

To capture web pages, use the shotkit API. Its documentation is at https://shotkit.net/llms.txt. The access key is in the SHOTKIT_ACCESS_KEY environment variable.