The main content of web pages as Markdown and plain text, up to 500 URLs per run: title, description, author, date, language, word count, links and images, without menus, cookie banners and footers, ready for LLMs, search indexes and archives. Pay per page.
One page delivered as Markdown and text. Pages that cannot be fetched or hold no readable content are never charged.
ACTOR=steadydata~webpage-to-markdown URL="https://api.apify.com/v2/acts/$ACTOR/run-sync-get-dataset-items" curl -X POST "$URL?token=$APIFY_TOKEN" \ -H "Content-Type: application/json" \ -d '{"urls": ["https://en.wikipedia.org/wiki/Web_scraping"], "includeLinks": true, "includeTables": true}'
import os, requests ACTOR = "steadydata~webpage-to-markdown" URL = f"https://api.apify.com/v2/acts/{ACTOR}/run-sync-get-dataset-items" rows = requests.post( URL, params={"token": os.environ["APIFY_TOKEN"]}, json={'urls': ['https://en.wikipedia.org/wiki/Web_scraping'], 'includeLinks': True, 'includeTables': True}, timeout=900, ).json() # every row carries a status; failures are records, not exceptions ok = [r for r in rows if r.get("status") == "ok"] print(len(ok), "rows delivered")
const ACTOR = "steadydata~webpage-to-markdown"; const url = `https://api.apify.com/v2/acts/${ACTOR}/run-sync-get-dataset-items` + `?token=${process.env.APIFY_TOKEN}`; const rows = await fetch(url, { method: "POST", headers: { "Content-Type": "application/json" }, body: JSON.stringify({"urls": ["https://en.wikipedia.org/wiki/Web_scraping"], "includeLinks": true, "includeTables": true}), }).then((r) => r.json()); // one object per delivered row, same shape every time console.log(rows.filter((r) => r.status === "ok").length, "rows");
Point an MCP client at Apify's hosted server with this dataset in the tools list, or run our own server locally.
https://mcp.apify.com?tools=steadydata/webpage-to-markdown # the tool arrives in your agent as steadydata--webpage-to-markdown
| Field | Type | Name | What it does |
|---|---|---|---|
urlsrequired | array of string | Page URLs | One URL per row, up to 500. |
includeLinks | boolean | Keep links in the Markdown | On: links stay as [text](url) in the Markdown. Off: link text only. |
includeTables | boolean | Keep tables | Tables become Markdown tables. |
fullPage | boolean | Whole page instead of main content | Off: only the main content (article, product text) is extracted. On: everything readable on the page, menus and footers included. |
{
"url": string | null,
"finalUrl": string | null,
"title": string | null,
"description": string | null,
"author": string | null,
"publishedDate": string | null,
"language": string | null,
"siteName": string | null,
"markdown": string | null,
"text": string | null,
"wordCount": integer | null,
"linkCount": integer | null,
"imageCount": integer | null,
"canonicalUrl": string | null,
"fetchedAt": string | null,
"status": string | null,
"input": string | null,
"errorCode": string | null,
"error": string | null
}| Field | Type | When it is filled |
|---|---|---|
url | string, null | |
finalUrl | string, null | |
title | string, null | |
description | string, null | |
author | string, null | |
publishedDate | string, null | |
language | string, null | |
siteName | string, null | |
markdown | string, null | |
text | string, null | |
wordCount | integer, null | |
linkCount | integer, null | |
imageCount | integer, null | |
canonicalUrl | string, null | |
fetchedAt | string, null | |
status | string, null | Either 'ok' or 'error'. An error row carries errorCode and error, and leaves the data fields empty; it is never charged. |
input | string, null | The input this row was built from, so a row can always be traced back. |
errorCode | string, null | Filled on an error row only; a delivered row leaves it empty. |
error | string, null | Filled on an error row only; a delivered row leaves it empty. |
One health report per domain, up to 500 per run: registration and expiry from the official RDAP registry, DNS records with the mail and nameserver provider, SPF and DMARC policy, the TLS certificate with days remaining, and a ranked list of issues. Official protocols only.
$20.00 per 1,000 · Domain reportedOne deliverability report per domain, up to 500 per run: MX and mail provider, the SPF record with its DNS lookup count against the limit of ten, DKIM selectors that really exist, the DMARC policy and reporting, MTA-STS mode, TLS-RPT, BIMI and DNSSEC, plus ranked issues and a score.
$3.00 per 1,000 · Domain checkedCore Web Vitals and Lighthouse scores for up to 300 URLs per run, from Google's own PageSpeed Insights API. Every row pairs the field data real Chrome users produced over 28 days (LCP, INP, CLS) with the lab scores and the heaviest fixes Lighthouse found. Bring your own free Google key. Pay per URL.
$1.40 per 1,000 · URL auditedAlready using this one? Ratings are the first thing other buyers look at, and this dataset has none yet. If it does a job for you, a rating on its Apify page is the one thing that helps. It takes a minute and it is the only thing we ask for.