PDF to Text in PHP
Overview
Extract text from PDF documents with the PDFCrowd PHP client. The client handles communication with the API, while conversions run on PDFCrowd's servers.
Installation
Install the PHP client with Composer, or see other installation options.
composer require pdfcrowd/pdfcrowd
Quick Start
The examples below use your API credentials. The examples use demo credentials, so you can try the API before getting your own API key. Update the input URLs and filenames as needed.
Convert a PDF URL
Extract text from a PDF URL and save it locally as invoice.txt:
<?php require 'vendor/autoload.php'; $client = new \Pdfcrowd\PdfToTextClient('demo', 'demo'); $client->convertUrlToFile('https://your-server.com/invoice.pdf', 'invoice.txt');
The URL must return a PDF and be reachable from PDFCrowd's servers.
The client throws Pdfcrowd\Error on conversion or validation errors.
See Handle Errors.
Convert a PDF File
Extract text from invoice.pdf and save the result locally as invoice.txt:
<?php require 'vendor/autoload.php'; $client = new \Pdfcrowd\PdfToTextClient('demo', 'demo'); $client->convertFileToFile('invoice.pdf', 'invoice.txt');
Configure a Conversion
Authentication
Pass your PDFCrowd username and API key to
PdfToTextClient.
Find your credentials on the API Keys page.
Choose an Input
Send a PDF URL, upload a local file, or pass PDF bytes or a readable binary stream.
| Input | Method |
|---|---|
| URL | convertUrlToFile() |
| Local file | convertFileToFile() |
| Bytes in memory | convertRawDataToFile() |
| Readable binary stream | convertStreamToFile() |
Add Conversion Settings
Set conversion options on the client before calling a conversion method.
For example, $client->setNoLayout(true); extracts text without preserving the PDF layout.
Common settings are listed below. See the method reference for all available options or browse PHP examples.
| Purpose | Methods |
|---|---|
| Pages to convert | setPrintPageRange() |
| Password-protected input | setPdfPassword() |
| Layout and paragraphs | setNoLayout(), setParagraphMode(), setLineSpacingThreshold() |
| Line endings and page breaks | setEol(), setPageBreakMode(), setCustomPageBreak() |
| Text cleanup | setRemoveHyphenation(), setRemoveEmptyLines() |
| Extract from a page area | setCropArea() |
By default, extraction preserves the PDF's layout with spacing. Enable setNoLayout() to extract text without preserving columns and positioning.
The PDF password unlocks the input document; it is separate from your PDFCrowd API key.
Handle the Result
Choose an Output
The result is UTF-8 plain text. A PDF must contain a text layer for extraction; recognizing text in a scanned document requires a separate OCR step.
The methods below use URL input; local files, bytes, and input streams have corresponding methods.
| Output | Method |
|---|---|
| Local file | convertUrlToFile() |
| String in memory | convertUrl() |
| Writable stream resource | convertUrlToStream() |
PHP strings can hold binary data. This example extracts text from a PDF,
receives the result as a string and saves it as invoice.txt:
<?php require 'vendor/autoload.php'; $client = new \Pdfcrowd\PdfToTextClient('demo', 'demo'); $result = $client->convertUrl('https://your-server.com/invoice.pdf'); if (file_put_contents('invoice.txt', $result) !== strlen($result)) { throw new \RuntimeException('Cannot save the conversion result'); }
When serving the result from a web application, use Content-Type: text/plain; charset=utf-8.
Handle Errors
The client throws a Pdfcrowd\Error
exception on conversion or validation errors.
This example extracts text from a PDF and logs any PDFCrowd error, including its
HTTP status and reason code:
<?php require 'vendor/autoload.php'; $client = new \Pdfcrowd\PdfToTextClient('demo', 'demo'); try { $client->convertUrlToFile( 'https://your-server.com/invoice.pdf', 'invoice.txt' ); } catch (\Pdfcrowd\Error $error) { error_log('PDFCrowd: ' . $error); error_log(sprintf('Status: %s; reason: %s', $error->getStatusCode(), $error->getReasonCode())); throw $error; }
Local PHP errors, such as a failure to read an input file or write the output, may need separate handling.
Pdfcrowd\Error provides these methods:
| Method | Returns |
|---|---|
getStatusCode() | The HTTP status code, when available. |
getReasonCode() | The reason code identifying the specific error, or -1 if unavailable. |
getMessage() | The error message. |
getDocumentationLink() | A link to relevant documentation, when available. |
(string) $error returns the complete error, including available status and reason codes.
Common Status Codes
| Status | What it means | What to do |
|---|---|---|
400 | Invalid input, settings, or conversion failure | Read the reason code and correct the input or settings. |
401 | Missing credentials or an inactive license | Check your username, API key, and license status. |
403 | Suspended service or no credits remaining | Check your account and available credits. |
413 | Upload exceeds the 300 MB limit | Reduce the upload size. |
429 | Request rate limit reached | Wait and reduce the rate of new requests. |
430 | Concurrent request limit reached | Allow active requests to finish before starting more. |
503 | Temporary network issue | Check your retry policy before submitting another request. |
See all status and reason codes for specific explanations and Limits and Retries for retry behavior.
Read Conversion Information
This information is available after a conversion and describes the client's last conversion.
| Method | Use |
|---|---|
getJobId() | Identify the conversion in logs and support requests. |
getDebugLogUrl() | The URL of the conversion debug log when logging is enabled. |
getPageCount() | Read the page count reported for the conversion. |
getOutputSize() | Read the text size in bytes. |
getConsumedCreditCount() | Read the credits consumed by the conversion. |
getRemainingCreditCount() | Read the remaining credit count reported with the conversion. |
Limits and Retries
Request rate and concurrency limits depend on your license. Control how quickly
your application submits conversions and how many it runs at once. A
429 response concerns request rate; a 430 response
concerns requests already in progress. The maximum upload size is 300 MB.
The PHP client automatically retries a request once when it receives HTTP
502 or 503. Use
setRetryCount()
to change that count, or set it to 0 to disable automatic retries.
Account for these retries when adding an application-level retry policy.
Troubleshooting
Inspect a Conversion
This example extracts text from a PDF with debug logging enabled and records the debug log URL when available. Use the input and settings from the conversion you are investigating.
<?php require 'vendor/autoload.php'; $client = new \Pdfcrowd\PdfToTextClient('demo', 'demo'); try { $client->setDebugLog(true); $client->convertUrlToFile( 'https://your-server.com/invoice.pdf', 'invoice.txt' ); } catch (\Pdfcrowd\Error $error) { error_log('PDFCrowd: ' . $error); error_log(sprintf('Status: %s; reason: %s', $error->getStatusCode(), $error->getReasonCode())); throw $error; } finally { if ($client->getDebugLogUrl()) { error_log('Debug log: ' . $client->getDebugLogUrl()); } }
The debug log contains conversion settings and processing details. You can also find logs in your conversion history. A local or connection failure may occur before a conversion log is available.
Common Problems
| Problem | Check |
|---|---|
| The PDF cannot be loaded | Check that the URL returns a PDF, rather than HTML or a login page, and is reachable from PDFCrowd's servers. For local files, check the path and read permissions. |
| The PDF requires a password | Set the input document's password with setPdfPassword(). |
| Text is missing | A scanned PDF may contain images without a text layer. Extracting text from scans requires a separate OCR step. |
| Spacing or reading order is unexpected | Try setNoLayout() and the paragraph settings. PDF text placement does not always encode a natural reading order. |
For help, contact support and include any available diagnostics, the client version, and enough detail to reproduce the problem.