HTML to PDF in Java
Overview
Convert web pages and HTML documents to PDF with the PDFCrowd Java client. The client handles communication with the API, while conversions run on PDFCrowd's servers.
Installation
Add the Java client to the <dependencies> section of your
pom.xml, or see
other installation options.
<dependency> <groupId>com.pdfcrowd</groupId> <artifactId>pdfcrowd</artifactId> <version>6.7.0</version> </dependency>
Quick Start
The examples below use your API credentials. The examples use demo credentials, so you can try the API before getting your own API key. Update the input URLs and filenames as needed.
Convert a URL
Convert a web page to PDF and save it locally as example.pdf:
import com.pdfcrowd.Pdfcrowd; class Example { public static void main(String[] args) throws Exception { Pdfcrowd.HtmlToPdfClient client = new Pdfcrowd.HtmlToPdfClient("demo", "demo"); client.setContentViewportWidth("balanced"); client.convertUrlToFile("https://example.com/", "example.pdf"); } }
The balanced viewport setting fits the HTML viewport to the print area.
The client throws Pdfcrowd.Error on conversion or validation errors.
See Handle Errors.
Convert an HTML File
Upload document.html and save the converted PDF locally as
document.pdf:
import com.pdfcrowd.Pdfcrowd; class Example { public static void main(String[] args) throws Exception { Pdfcrowd.HtmlToPdfClient client = new Pdfcrowd.HtmlToPdfClient("demo", "demo"); client.setContentViewportWidth("balanced"); client.convertFileToFile("document.html", "document.pdf"); } }
For HTML that references local images or stylesheets, see Include Local Assets. To pass an HTML string, see Send HTML Content.
Configure a Conversion
Authentication
Pass your PDFCrowd username and API key to
HtmlToPdfClient.
Find your credentials on the API Keys page.
To load a protected source website, configure its website credentials, cookies, or HTTP headers separately.
Choose an Input
| Input | Method |
|---|---|
| Web page URL | convertUrlToFile() |
| HTML string | convertStringToFile() |
| Local HTML file or archive | convertFileToFile() |
URLs must be reachable from PDFCrowd's servers. For a page on
localhost, send its HTML content or upload a file.
Referenced assets must also be reachable from PDFCrowd's servers or included
in an archive.
Send HTML Content
Convert an HTML string to PDF and save it as report.pdf:
import com.pdfcrowd.Pdfcrowd; class Example { public static void main(String[] args) throws Exception { Pdfcrowd.HtmlToPdfClient client = new Pdfcrowd.HtmlToPdfClient("demo", "demo"); String html = "<!doctype html>\n" + "<html><body><h1>Monthly report</h1></body></html>"; client.setContentViewportWidth("balanced"); client.convertStringToFile(html, "report.pdf"); } }
Replace html with the output of your template renderer.
For relative resource URLs, add a
<base href="https://example.com/"> element to the HTML head,
or use absolute URLs. PDFCrowd can also combine an HTML template with structured
data; see HTML templates.
Include Local Assets
Package the HTML and its local images, CSS, and JavaScript in a
.zip, .tar.gz, or .tar.bz2 archive.
Preserve the relative paths used by the HTML.
Convert the HTML document inside document.zip and save the PDF
as output.pdf:
import com.pdfcrowd.Pdfcrowd; class Example { public static void main(String[] args) throws Exception { Pdfcrowd.HtmlToPdfClient client = new Pdfcrowd.HtmlToPdfClient("demo", "demo"); client.setContentViewportWidth("balanced"); client.convertFileToFile("document.zip", "output.pdf"); } }
The API automatically converts the first HTML file it finds in the archive.
If the archive contains multiple HTML files, use
setZipMainFilename() to choose which one to convert.
Add Conversion Settings
Set conversion options on the client before calling a conversion method.
For example, client.setPageSize("A4"); selects A4 paper.
Common settings are listed below. See the method reference for all available options, browse Java examples, or try settings in the API Playground.
| Purpose | Methods |
|---|---|
| Page size and margins | setPageSize(), setPageMargins(), setNoMargins(), setOrientation() |
| Headers and footers | setHeaderHtml(), setFooterHtml(), setHeaderHeight(), setFooterHeight() |
| Layout and scaling | setContentViewportWidth(), setScaleFactor() |
| Print styles and custom CSS | setUsePrintMedia(), setCustomCss() |
| JavaScript and readiness | setCustomJavascript(), setWaitForElement(), setJavascriptDelay() |
| Convert a specific element | setElementToConvert() |
Handle the Result
Choose an Output
The methods below use URL input; HTML strings and files have corresponding methods.
| Output | Method |
|---|---|
| Local file | convertUrlToFile() |
| Byte array | convertUrl() |
OutputStream | convertUrlToStream() |
File-output methods overwrite an existing destination file.
This example converts a web page, receives the result as a
byte[] and saves it as example.pdf:
import com.pdfcrowd.Pdfcrowd; import java.nio.file.Files; import java.nio.file.Path; class Example { public static void main(String[] args) throws Exception { Pdfcrowd.HtmlToPdfClient client = new Pdfcrowd.HtmlToPdfClient("demo", "demo"); byte[] result = client.convertUrl( "https://example.com/" ); Files.write(Path.of("example.pdf"), result); } }
When serving the PDF from a web application, use Content-Type: application/pdf.
Handle Errors
The client throws a Pdfcrowd.Error
exception on conversion or validation errors.
This example converts a web page and logs any PDFCrowd error, including its
HTTP status and reason code:
import com.pdfcrowd.Pdfcrowd; import java.util.logging.Level; import java.util.logging.Logger; class Example { private static final Logger logger = Logger.getLogger(Example.class.getName()); public static void main(String[] args) throws Exception { Pdfcrowd.HtmlToPdfClient client = new Pdfcrowd.HtmlToPdfClient("demo", "demo"); try { client.convertUrlToFile( "https://example.com/", "example.pdf" ); } catch (Pdfcrowd.Error error) { logger.log(Level.SEVERE, "PDFCrowd conversion failed (status " + error.getStatusCode() + ", reason " + error.getReasonCode() + ")", error); throw error; } } }
Local Java errors, such as a failure to read an input file or write the output, may need separate handling.
Pdfcrowd.Error provides these methods:
| Method | Returns |
|---|---|
getStatusCode() | The HTTP status code, when available. |
getReasonCode() | The reason code identifying the specific error, or -1 if unavailable. |
getMessage() | The error message. |
getDocumentationLink() | A link to relevant documentation, when available. |
Common Status Codes
| Status | What it means | What to do |
|---|---|---|
400 | Invalid input, settings, or conversion failure | Read the reason code and correct the input or settings. |
401 | Missing credentials or an inactive license | Check your username, API key, and license status. |
403 | Suspended service or no credits remaining | Check your account and available credits. |
413 | Upload exceeds the 300 MB limit | Reduce the upload size. |
429 | Request rate limit reached | Wait and reduce the rate of new requests. |
430 | Concurrent request limit reached | Allow active requests to finish before starting more. |
503 | Temporary network issue | Check your retry policy before submitting another request. |
See all status and reason codes for specific explanations and Limits and Retries for retry behavior.
Read Conversion Information
This information is available after a conversion and describes the client's last conversion.
| Method | Use |
|---|---|
getJobId() | Identify the conversion in logs and support requests. |
getDebugLogUrl() | The URL of the conversion debug log when logging is enabled. |
getPageCount() | Read the number of pages in the PDF. |
getOutputSize() | Read the PDF size in bytes. |
getConsumedCreditCount() | Read the credits consumed by the conversion. |
getRemainingCreditCount() | Read the remaining credit count reported with the conversion. |
Limits and Retries
Request rate and concurrency limits depend on your license. Control how quickly
your application submits conversions and how many it runs at once. A
429 response concerns request rate; a 430 response
concerns requests already in progress. The maximum upload size is 300 MB.
The Java client automatically retries a request once when it receives HTTP
502 or 503. Use
setRetryCount()
to change that count, or set it to 0 to disable automatic retries.
Account for these retries when adding an application-level retry policy.
PDFCrowd stops a conversion that exceeds 60 seconds of processing time.
Reason code 323 identifies this HTML conversion timeout.
Check for slow resource downloads or long-running JavaScript. Allowing your
application to wait longer does not extend the server's processing limit.
Troubleshooting
Inspect a Conversion
This example converts a web page with debug logging enabled and logs the debug log URL when available. Use the input and settings from the conversion you are investigating.
import com.pdfcrowd.Pdfcrowd; import java.util.logging.Level; import java.util.logging.Logger; class Example { private static final Logger logger = Logger.getLogger(Example.class.getName()); public static void main(String[] args) throws Exception { Pdfcrowd.HtmlToPdfClient client = new Pdfcrowd.HtmlToPdfClient("demo", "demo"); try { client.setDebugLog(true); client.convertUrlToFile( "https://example.com/", "example.pdf" ); } catch (Pdfcrowd.Error error) { logger.log(Level.SEVERE, "PDFCrowd conversion failed (status " + error.getStatusCode() + ", reason " + error.getReasonCode() + ")", error); throw error; } finally { String debugLogUrl = client.getDebugLogUrl(); if (debugLogUrl != null && !debugLogUrl.isEmpty()) { logger.info("Debug log: " + debugLogUrl); } } } }
The debug log includes resource-loading details, timeouts, and browser console messages. You can also find logs in your conversion history. A local or connection failure may occur before a conversion log is available.
Common Problems
| Problem | Check |
|---|---|
| Images or styles are missing | Check resource URLs and authentication. See Include Local Assets for local resources. |
| A login page appears in the output | Check the source website's authentication. |
| Dynamic content is missing | Use setWaitForElement() or setJavascriptDelay(), and inspect the debug log. |
| The layout has the wrong width | Adjust setContentViewportWidth(). |
| The output needs print-specific styling | Enable setUsePrintMedia() if the page supplies print styles, or add setCustomCss(). |
| The application times out before receiving a result | Check client and application timeouts. Allow for uploading the input, conversion, and downloading the result. |
For help, contact support and include any available diagnostics, the client version, and enough detail to reproduce the problem.