PDF to Text HTTP API

Overview

Extract text from PDF documents with a direct HTTP request from any language or platform. For language-specific client libraries, see the SDK guides.

RequestDetails
Method and endpoint POST https://api.pdfcrowd.com/convert/24.04/
Authentication HTTP Basic: your PDFCrowd username and API key
Request body Form fields; use multipart form data for file uploads
Successful response 200 OK with the extracted text in the response body

Quick Start

Choose an input below. You can run the examples below with the displayed API credentials. Update the input URLs and filenames as needed.

Each example uses cURL and saves the returned text to a local file. If a conversion fails, inspect the response for error details.

Convert a PDF URL

Send the PDF's address in the url field. This example extracts its text and saves the result in invoice.txt:

curl -f -s -S \
  -u 'demo:demo' \
  -o invoice.txt \
  -F url=https://your-server.com/invoice.pdf \
  -F input_format=pdf \
  -F output_format=txt \
  https://api.pdfcrowd.com/convert/24.04/

PDFCrowd downloads the PDF and returns the result in the same request. The default output preserves the PDF's layout, including column spacing.

Convert a PDF File

Upload an existing PDF in the file field. Replace the path with your local file; this example saves the extracted text in invoice.txt:

curl -f -s -S \
  -u 'demo:demo' \
  -o invoice.txt \
  -F file=@invoice.pdf \
  -F input_format=pdf \
  -F output_format=txt \
  https://api.pdfcrowd.com/convert/24.04/

Open the saved text file or read it in your application. For page selection, layout and text cleanup, see Add Conversion Settings.

cURL example options

Use a Bash-compatible shell.

Basic options

  • -u — Supply your username and API key.
  • -F — Send a form field.
  • -o — Save the response to a file.
  • -f — Report HTTP errors as command failures.
  • -s -S — Hide the progress meter while keeping error messages.

File input

  • file=@document.pdf — Upload a local file.

Diagnostics

  • --fail-with-body — Use instead of -f to retain the error response body.
  • -D — Save response headers to a file.
  • -w — Print response information; the diagnostic example prints the HTTP status.

Build a Request

Authentication

Use your PDFCrowd username as the HTTP Basic username and your API key as the password. Configure these credentials using your HTTP client's Basic authentication option.

Find and manage your credentials on the API Keys page.

Request Format

Send the input and conversion settings as form fields in a POST request. JSON request bodies are not supported. Your HTTP client handles form encoding and the appropriate headers.

Use the versioned endpoint: https://api.pdfcrowd.com/convert/24.04/. Keep the version explicit in your integration and review API versioning before changing it.

Errors are returned as plain text by default; add ?errfmt=json to the endpoint to receive JSON errors. This changes only the error format; a successful response still contains the extracted text. See Handle the Response for response handling.

Choose an Input

Set input_format=pdf and output_format=txt. Supply one of the following inputs:

FieldWhat to Send
url An http:// or https:// address that returns a PDF and is reachable from PDFCrowd's servers.
file The PDF's contents uploaded as a multipart file part.

Upload the file's contents, not its filename or local path as a text field. If the PDF is already in memory, use your HTTP client's file-upload option to send those bytes in the file part. PDFCrowd cannot fetch a PDF from your computer's localhost; upload its contents instead.

Add Conversion Settings

The default output preserves the PDF's layout. Set no_layout to true for reading-order output without column spacing or positioning. Choose the mode suited to your document; neither guarantees reconstruction of semantic order or tables.

Send conversion settings as additional form fields alongside your input. Common options include:

PurposeOptions
Page selection print_page_range
Text layout no_layout
Line endings and page separators eol, page_break_mode, custom_page_break
Paragraph detection paragraph_mode, line_spacing_threshold
Text cleanup remove_hyphenation, remove_empty_lines
Extract text from a page region crop_area_x, crop_area_y, crop_area_width, crop_area_height
Password-protected input pdf_password

The parameter reference lists all settings, accepted values, defaults, and constraints. For complete conversion requests, see the HTTP examples.

Handle the Response

Successful Conversion

A successful conversion returns 200 OK with Content-Type: text/plain and the extracted text in the response body. Save it to a text file or process it in your application.

An illustrative response is:

HTTP/1.1 200 OK
Content-Type: text/plain
x-pdfcrowd-job-id: example-job-id

[Extracted text]

Check the HTTP status before using the body. Both extracted text and the default error response use text/plain, so the content type alone does not indicate success. Setting errfmt=json changes errors to JSON; successful conversions still return plain text.

This API extracts the PDF's existing text layer. Image-only scans need a separate OCR step before their text can be extracted.

Errors

An unsuccessful request returns an error HTTP status. PDFCrowd error responses also include a reason code that identifies the specific problem.

By default, the error body is plain text in this format:

<status_code>.<reason_code> - <message>

For structured errors, append ?errfmt=json to the endpoint:

https://api.pdfcrowd.com/convert/24.04/?errfmt=json

An example JSON error body for missing conversion input is:

{
  "status_code": 400,
  "reason_code": 325,
  "message": "There is no input specified to be converted."
}

PDFCrowd JSON errors use Content-Type: application/json. A successful request still returns plain text when errfmt=json is set. If a failed response has a different content type, retain its body and status for diagnosis instead of assuming it is JSON.

Common Status Codes

StatusWhat it meansWhat to do
400 Invalid request or conversion failure Read the reason code and correct the input or settings.
401 Missing credentials or an inactive license Check your username, API key, and license status.
403 Suspended service or no credits remaining Check your account and available credits.
413 Upload exceeds the 300 MB limit Reduce the upload size.
429 Request rate limit reached Wait and reduce the rate of new requests.
430 Concurrent request limit reached Allow active requests to finish before starting more.
503 Temporary network issue Retry after a delay.

See all status and reason codes for specific error explanations. For request limits and retry guidance, see Limits and Retries.

Response Headers

Use these headers, when present, to record conversion results and diagnose problems:

Header Use
x-pdfcrowd-job-id Identify the conversion in logs and support requests.
x-pdfcrowd-reason-code Read the error reason code; 0 indicates success.
x-pdfcrowd-debug-log Open the debug log when debug logging is enabled.
x-pdfcrowd-consumed-credits Record credits consumed by this conversion.
x-pdfcrowd-remaining-credits Monitor the remaining account balance.
x-pdfcrowd-output-size Read the output size in bytes.

Limits and Retries

Request rate and concurrency limits depend on your license. Control how quickly you submit conversions and how many you run at once. A 429 response concerns request rate; a 430 response concerns requests already in progress.

The maximum upload size is 300 MB.

For temporary failures, use a bounded number of retries with increasing delays. Correct invalid input, authentication, or account problems before retrying those requests.

If a conversion exceeds 60 seconds of processing time, PDFCrowd stops it and returns an error response.

Troubleshooting

Inspect a Request

Capture the HTTP status, response headers, and response body when diagnosing a request. Set debug_log to true to enable a conversion debug log, and add errfmt=json to the endpoint's query string for structured error details. Use the input and conversion settings from the request you're investigating.

curl --fail-with-body \
  -u 'demo:demo' \
  -D response.headers \
  -o response.body \
  -w 'HTTP %{http_code}\n' \
  -F url=https://your-server.com/invoice.pdf \
  -F input_format=pdf \
  -F output_format=txt \
  -F debug_log=true \
  'https://api.pdfcrowd.com/convert/24.04/?errfmt=json'

Check the HTTP status before interpreting the body. A successful response contains extracted text; an unsuccessful response should be inspected for error details. If no HTTP response arrives, check the error reported by your HTTP client.

The x-pdfcrowd-debug-log header links to diagnostic information about the conversion. You can also find logs in your conversion history.

Common Problems

ProblemCheck
No input or no request data Send form fields rather than JSON. Supply url or upload the PDF's contents in a multipart file part named file.
The PDF cannot be loaded Check that the URL returns the PDF rather than a login page and is reachable from PDFCrowd's servers. For a local file, check that your client uploads its contents.
The PDF cannot be opened Check that the input is a valid PDF. If it is password-protected, supply pdf_password. This is separate from your PDFCrowd API key.
The result is empty or some text is missing Check that the PDF has a text layer; image-only scans require a separate OCR step. Check print_page_range and the crop area settings for excluded pages or content.
Spacing or reading order is unexpected Try no_layout=true for output without column spacing. Text order also depends on the PDF's internal structure; extraction does not guarantee reconstruction of tables or semantic reading order.
Page separators are missing The default page_break_mode is none. Choose default for form-feed characters, or custom with custom_page_break.
The text contains unwanted hyphens or empty lines Use remove_hyphenation for line-end hyphens and remove_empty_lines for blank lines. Check that the result suits your document and downstream use.
The HTTP client times out before receiving the response Check its timeout settings and allow enough time for conversion, uploading the input, and downloading the result.

For help, contact support and include any available diagnostics and enough detail for us to reproduce the problem.