PDF to Text HTTP API
Overview
Extract text from PDF documents with a direct HTTP request from any language or platform. For language-specific client libraries, see the SDK guides.
| Request | Details |
|---|---|
| Method and endpoint | POST https://api.pdfcrowd.com/convert/24.04/ |
| Authentication | HTTP Basic: your PDFCrowd username and API key |
| Request body | Form fields; use multipart form data for file uploads |
| Successful response | 200 OK with the extracted text in the response body |
Quick Start
Choose an input below. You can run the examples below with the displayed API credentials. Update the input URLs and filenames as needed.
Each example uses cURL and saves the returned text to a local file. If a conversion fails, inspect the response for error details.
Convert a PDF URL
Send the PDF's address in the url field. This example extracts
its text and saves the result in invoice.txt:
curl -f -s -S \ -u 'demo:demo' \ -o invoice.txt \ -F url=https://your-server.com/invoice.pdf \ -F input_format=pdf \ -F output_format=txt \ https://api.pdfcrowd.com/convert/24.04/
PDFCrowd downloads the PDF and returns the result in the same request. The default output preserves the PDF's layout, including column spacing.
Convert a PDF File
Upload an existing PDF in the file field. Replace the path with
your local file; this example saves the extracted text in invoice.txt:
curl -f -s -S \ -u 'demo:demo' \ -o invoice.txt \ -F file=@invoice.pdf \ -F input_format=pdf \ -F output_format=txt \ https://api.pdfcrowd.com/convert/24.04/
Open the saved text file or read it in your application. For page selection, layout and text cleanup, see Add Conversion Settings.
cURL example options
Use a Bash-compatible shell.
Basic options
-u— Supply your username and API key.-F— Send a form field.-o— Save the response to a file.-f— Report HTTP errors as command failures.-s -S— Hide the progress meter while keeping error messages.
File input
file=@document.pdf— Upload a local file.
Diagnostics
-
--fail-with-body— Use instead of-fto retain the error response body. -D— Save response headers to a file.-
-w— Print response information; the diagnostic example prints the HTTP status.
Build a Request
Authentication
Use your PDFCrowd username as the HTTP Basic username and your API key as the password. Configure these credentials using your HTTP client's Basic authentication option.
Find and manage your credentials on the API Keys page.
Request Format
Send the input and conversion settings as form fields in a POST request. JSON request bodies are not supported. Your HTTP client handles form encoding and the appropriate headers.
Use the versioned endpoint: https://api.pdfcrowd.com/convert/24.04/.
Keep the version explicit in your integration and review
API versioning before changing it.
Errors are returned as plain text by default; add ?errfmt=json
to the endpoint to receive JSON errors. This changes only the error format;
a successful response still contains the extracted text.
See Handle the Response for response handling.
Choose an Input
Set input_format=pdf and output_format=txt.
Supply one of the following inputs:
| Field | What to Send |
|---|---|
url |
An http:// or https:// address that returns a PDF and is reachable from PDFCrowd's servers. |
file |
The PDF's contents uploaded as a multipart file part. |
Upload the file's contents, not its filename or local path as a text field.
If the PDF is already in memory, use your HTTP client's file-upload option
to send those bytes in the file part.
PDFCrowd cannot fetch a PDF from your computer's localhost;
upload its contents instead.
Add Conversion Settings
The default output preserves the PDF's layout. Set
no_layout
to true for reading-order output without column spacing or positioning.
Choose the mode suited to your document; neither guarantees reconstruction
of semantic order or tables.
Send conversion settings as additional form fields alongside your input. Common options include:
| Purpose | Options |
|---|---|
| Page selection | print_page_range |
| Text layout | no_layout |
| Line endings and page separators |
eol,
page_break_mode,
custom_page_break
|
| Paragraph detection |
paragraph_mode,
line_spacing_threshold
|
| Text cleanup |
remove_hyphenation,
remove_empty_lines
|
| Extract text from a page region |
crop_area_x,
crop_area_y,
crop_area_width,
crop_area_height
|
| Password-protected input | pdf_password |
The parameter reference lists all settings, accepted values, defaults, and constraints. For complete conversion requests, see the HTTP examples.
Handle the Response
Successful Conversion
A successful conversion returns 200 OK with
Content-Type: text/plain and the extracted text in the response
body. Save it to a text file or process it in your application.
An illustrative response is:
HTTP/1.1 200 OK Content-Type: text/plain x-pdfcrowd-job-id: example-job-id [Extracted text]
Check the HTTP status before using the body. Both extracted text and the
default error response use text/plain, so the content type alone
does not indicate success. Setting errfmt=json changes errors to
JSON; successful conversions still return plain text.
This API extracts the PDF's existing text layer. Image-only scans need a separate OCR step before their text can be extracted.
Errors
An unsuccessful request returns an error HTTP status. PDFCrowd error responses also include a reason code that identifies the specific problem.
By default, the error body is plain text in this format:
<status_code>.<reason_code> - <message>
For structured errors, append ?errfmt=json to the endpoint:
https://api.pdfcrowd.com/convert/24.04/?errfmt=json
An example JSON error body for missing conversion input is:
{
"status_code": 400,
"reason_code": 325,
"message": "There is no input specified to be converted."
}
PDFCrowd JSON errors use Content-Type: application/json. A
successful request still returns plain text
when errfmt=json is set. If a failed response has a different
content type, retain its body and status for diagnosis instead of assuming
it is JSON.
Common Status Codes
| Status | What it means | What to do |
|---|---|---|
400 |
Invalid request or conversion failure | Read the reason code and correct the input or settings. |
401 |
Missing credentials or an inactive license | Check your username, API key, and license status. |
403 |
Suspended service or no credits remaining | Check your account and available credits. |
413 |
Upload exceeds the 300 MB limit | Reduce the upload size. |
429 |
Request rate limit reached | Wait and reduce the rate of new requests. |
430 |
Concurrent request limit reached | Allow active requests to finish before starting more. |
503 |
Temporary network issue | Retry after a delay. |
See all status and reason codes for specific error explanations. For request limits and retry guidance, see Limits and Retries.
Response Headers
Use these headers, when present, to record conversion results and diagnose problems:
| Header | Use |
|---|---|
x-pdfcrowd-job-id |
Identify the conversion in logs and support requests. |
x-pdfcrowd-reason-code |
Read the error reason code; 0 indicates success. |
x-pdfcrowd-debug-log |
Open the debug log when debug logging is enabled. |
x-pdfcrowd-consumed-credits |
Record credits consumed by this conversion. |
x-pdfcrowd-remaining-credits |
Monitor the remaining account balance. |
x-pdfcrowd-output-size |
Read the output size in bytes. |
Limits and Retries
Request rate and concurrency limits depend on your license. Control how
quickly you submit conversions and how many you run at once. A 429
response concerns request rate; a 430 response concerns requests
already in progress.
The maximum upload size is 300 MB.
For temporary failures, use a bounded number of retries with increasing delays. Correct invalid input, authentication, or account problems before retrying those requests.
If a conversion exceeds 60 seconds of processing time, PDFCrowd stops it and returns an error response.
Troubleshooting
Inspect a Request
Capture the HTTP status, response headers, and response body when diagnosing
a request. Set debug_log to true to enable a
conversion debug log, and add errfmt=json to the endpoint's
query string for structured error details. Use the input and conversion
settings from the request you're investigating.
curl --fail-with-body \ -u 'demo:demo' \ -D response.headers \ -o response.body \ -w 'HTTP %{http_code}\n' \ -F url=https://your-server.com/invoice.pdf \ -F input_format=pdf \ -F output_format=txt \ -F debug_log=true \ 'https://api.pdfcrowd.com/convert/24.04/?errfmt=json'
Check the HTTP status before interpreting the body. A successful response contains extracted text; an unsuccessful response should be inspected for error details. If no HTTP response arrives, check the error reported by your HTTP client.
The x-pdfcrowd-debug-log header links to diagnostic information
about the conversion. You can also find logs in your
conversion history.
Common Problems
| Problem | Check |
|---|---|
| No input or no request data | Send form fields rather than JSON. Supply url or upload the PDF's contents in a multipart file part named file. |
| The PDF cannot be loaded | Check that the URL returns the PDF rather than a login page and is reachable from PDFCrowd's servers. For a local file, check that your client uploads its contents. |
| The PDF cannot be opened |
Check that the input is a valid PDF. If it is password-protected, supply
pdf_password.
This is separate from your PDFCrowd API key.
|
| The result is empty or some text is missing |
Check that the PDF has a text layer; image-only scans require a separate OCR step.
Check print_page_range
and the crop area settings
for excluded pages or content.
|
| Spacing or reading order is unexpected |
Try no_layout=true
for output without column spacing. Text order also depends on the PDF's
internal structure; extraction does not guarantee reconstruction of tables
or semantic reading order.
|
| Page separators are missing |
The default page_break_mode
is none. Choose default for form-feed characters,
or custom with
custom_page_break.
|
| The text contains unwanted hyphens or empty lines |
Use remove_hyphenation
for line-end hyphens and
remove_empty_lines
for blank lines. Check that the result suits your document and downstream use.
|
| The HTTP client times out before receiving the response | Check its timeout settings and allow enough time for conversion, uploading the input, and downloading the result. |
For help, contact support and include any available diagnostics and enough detail for us to reproduce the problem.