Text extraction (OCR) from images by URL, up to 200 per run: every line with its confidence and position, the full text joined, image size and format, for screenshots, scans, receipts, menus and product labels, without a browser or an API key. Pay per image with text.
One image delivered with its text. Images without readable text, unreachable files and non-images are never charged.
ACTOR=steadydata~image-ocr-text URL="https://api.apify.com/v2/acts/$ACTOR/run-sync-get-dataset-items" curl -X POST "$URL?token=$APIFY_TOKEN" \ -H "Content-Type: application/json" \ -d '{"urls": ["https://raw.githubusercontent.com/madmaze/pytesseract/master/tests/data/test.png"], "minConfidence": "0.5"}'
import os, requests ACTOR = "steadydata~image-ocr-text" URL = f"https://api.apify.com/v2/acts/{ACTOR}/run-sync-get-dataset-items" rows = requests.post( URL, params={"token": os.environ["APIFY_TOKEN"]}, json={'urls': ['https://raw.githubusercontent.com/madmaze/pytesseract/master/tests/data/test.png'], 'minConfidence': '0.5'}, timeout=900, ).json() # every row carries a status; failures are records, not exceptions ok = [r for r in rows if r.get("status") == "ok"] print(len(ok), "rows delivered")
const ACTOR = "steadydata~image-ocr-text"; const url = `https://api.apify.com/v2/acts/${ACTOR}/run-sync-get-dataset-items` + `?token=${process.env.APIFY_TOKEN}`; const rows = await fetch(url, { method: "POST", headers: { "Content-Type": "application/json" }, body: JSON.stringify({"urls": ["https://raw.githubusercontent.com/madmaze/pytesseract/master/tests/data/test.png"], "minConfidence": "0.5"}), }).then((r) => r.json()); // one object per delivered row, same shape every time console.log(rows.filter((r) => r.status === "ok").length, "rows");
Point an MCP client at Apify's hosted server with this dataset in the tools list, or run our own server locally. The field explanations in the table below travel with the schema, so your agent reads them straight from the platform instead of from this page.
https://mcp.apify.com?tools=steadydata/image-ocr-text # the tool arrives in your agent as steadydata--image-ocr-text
| Field | Type | Name | What it does |
|---|---|---|---|
urlsrequired | array of string | Image URLs | One URL per row, up to 200: PNG, JPEG, WebP, GIF, BMP or TIFF. |
minConfidence | string | Minimum confidence per line | Lines below this OCR confidence (0 to 1) are dropped; 0.5 keeps almost everything, 0.8 only clean text. |
{
"url": string | null,
"text": string | null,
"lines": array | null,
"lineCount": integer | null,
"avgConfidence": number | null,
"boxes": array | null,
"width": integer | null,
"height": integer | null,
"format": string | null,
"fileBytes": integer | null,
"processedMs": integer | null,
"status": string | null,
"input": string | null,
"errorCode": string | null,
"error": string | null
}| Field | Type | When it is filled |
|---|---|---|
url | string, null | |
text | string, null | |
lines | array, null | |
lineCount | integer, null | |
avgConfidence | number, null | |
boxes | array, null | |
width | integer, null | |
height | integer, null | |
format | string, null | |
fileBytes | integer, null | |
processedMs | integer, null | |
status | string, null | Either 'ok' or 'error'. An error row carries errorCode and error, and leaves the data fields empty; it is never charged. |
input | string, null | The input this row was built from, so a row can always be traced back. |
errorCode | string, null | Filled on an error row only; a delivered row leaves it empty. |
error | string, null | Filled on an error row only; a delivered row leaves it empty. |
One health report per domain, up to 500 per run: registration and expiry from the official RDAP registry, DNS records with the mail and nameserver provider, SPF and DMARC policy, the TLS certificate with days remaining, and a ranked list of issues. Official protocols only.
$20.00 per 1,000 · Domain reportedOne deliverability report per domain, up to 500 per run: MX and mail provider, the SPF record with its DNS lookup count against the limit of ten, DKIM selectors that really exist, the DMARC policy and reporting, MTA-STS mode, TLS-RPT, BIMI and DNSSEC, plus ranked issues and a score.
$3.00 per 1,000 · Domain checkedCore Web Vitals and Lighthouse scores for up to 300 URLs per run, from Google's own PageSpeed Insights API. Every row pairs the field data real Chrome users produced over 28 days (LCP, INP, CLS) with the lab scores and the heaviest fixes Lighthouse found. Bring your own free Google key. Pay per URL.
$1.40 per 1,000 · URL auditedAlready using this one? Ratings are the first thing other buyers look at, and this dataset has none yet. If it does a job for you, a rating on its Apify page is the one thing that helps. It takes a minute and it is the only thing we ask for.