Plain text extraction from PDF files by URL, up to 200 files per run: one row per page (or one per file) with the text, page count, title, author, dates, producer and file size, plus whether the page has a text layer at all. Pay per page delivered.
One page of text delivered (or one file when pages are merged). Files that cannot be fetched or parsed are never charged.
ACTOR=steadydata~pdf-text-extractor URL="https://api.apify.com/v2/acts/$ACTOR/run-sync-get-dataset-items" curl -X POST "$URL?token=$APIFY_TOKEN" \ -H "Content-Type: application/json" \ -d '{"urls": ["https://arxiv.org/pdf/1706.03762"], "mergePages": false, "maxPagesPerFile": 200}'
import os, requests ACTOR = "steadydata~pdf-text-extractor" URL = f"https://api.apify.com/v2/acts/{ACTOR}/run-sync-get-dataset-items" rows = requests.post( URL, params={"token": os.environ["APIFY_TOKEN"]}, json={'urls': ['https://arxiv.org/pdf/1706.03762'], 'mergePages': False, 'maxPagesPerFile': 200}, timeout=900, ).json() # every row carries a status; failures are records, not exceptions ok = [r for r in rows if r.get("status") == "ok"] print(len(ok), "rows delivered")
const ACTOR = "steadydata~pdf-text-extractor"; const url = `https://api.apify.com/v2/acts/${ACTOR}/run-sync-get-dataset-items` + `?token=${process.env.APIFY_TOKEN}`; const rows = await fetch(url, { method: "POST", headers: { "Content-Type": "application/json" }, body: JSON.stringify({"urls": ["https://arxiv.org/pdf/1706.03762"], "mergePages": false, "maxPagesPerFile": 200}), }).then((r) => r.json()); // one object per delivered row, same shape every time console.log(rows.filter((r) => r.status === "ok").length, "rows");
Point an MCP client at Apify's hosted server with this dataset in the tools list, or run our own server locally.
https://mcp.apify.com?tools=steadydata/pdf-text-extractor # the tool arrives in your agent as steadydata--pdf-text-extractor
| Field | Type | Name | What it does |
|---|---|---|---|
urlsrequired | array of string | PDF URLs | One URL per row, up to 200. Direct links to PDF files. |
mergePages | boolean | One row per file instead of per page | Off: one row per page. On: one row per file with all pages joined; that row counts as one delivered page. |
maxPagesPerFile | integer | Max pages per file | Cost ceiling per file, from the first page. |
{
"url": string | null,
"fileName": string | null,
"page": integer | null,
"pageCount": integer | null,
"text": string | null,
"characters": integer | null,
"hasTextLayer": boolean | null,
"title": string | null,
"author": string | null,
"subject": string | null,
"createdAt": string | null,
"modifiedAt": string | null,
"producer": string | null,
"fileBytes": integer | null,
"encrypted": boolean | null,
"status": string | null,
"input": string | null,
"errorCode": string | null,
"error": string | null
}| Field | Type | When it is filled |
|---|---|---|
url | string, null | |
fileName | string, null | |
page | integer, null | |
pageCount | integer, null | |
text | string, null | |
characters | integer, null | |
hasTextLayer | boolean, null | |
title | string, null | |
author | string, null | |
subject | string, null | |
createdAt | string, null | |
modifiedAt | string, null | |
producer | string, null | |
fileBytes | integer, null | |
encrypted | boolean, null | |
status | string, null | Either 'ok' or 'error'. An error row carries errorCode and error, and leaves the data fields empty; it is never charged. |
input | string, null | The input this row was built from, so a row can always be traced back. |
errorCode | string, null | Filled on an error row only; a delivered row leaves it empty. |
error | string, null | Filled on an error row only; a delivered row leaves it empty. |
One health report per domain, up to 500 per run: registration and expiry from the official RDAP registry, DNS records with the mail and nameserver provider, SPF and DMARC policy, the TLS certificate with days remaining, and a ranked list of issues. Official protocols only.
$20.00 per 1,000 · Domain reportedOne deliverability report per domain, up to 500 per run: MX and mail provider, the SPF record with its DNS lookup count against the limit of ten, DKIM selectors that really exist, the DMARC policy and reporting, MTA-STS mode, TLS-RPT, BIMI and DNSSEC, plus ranked issues and a score.
$3.00 per 1,000 · Domain checkedCore Web Vitals and Lighthouse scores for up to 300 URLs per run, from Google's own PageSpeed Insights API. Every row pairs the field data real Chrome users produced over 28 days (LCP, INP, CLS) with the lab scores and the heaviest fixes Lighthouse found. Bring your own free Google key. Pay per URL.
$1.40 per 1,000 · URL auditedAlready using this one? Ratings are the first thing other buyers look at, and this dataset has none yet. If it does a job for you, a rating on its Apify page is the one thing that helps. It takes a minute and it is the only thing we ask for.