PDF to HTML in Java

Overview

Convert PDF documents to HTML with the PDFCrowd Java client. The client handles communication with the API, while conversions run on PDFCrowd's servers.

Installation

Add the Java client to the <dependencies> section of your pom.xml, or see other installation options.

<dependency>
  <groupId>com.pdfcrowd</groupId>
  <artifactId>pdfcrowd</artifactId>
  <version>6.7.0</version>
</dependency>

Quick Start

Update the input URLs and filenames as needed.

Convert a PDF URL

Convert a PDF URL to HTML and save it locally as invoice.html:

import com.pdfcrowd.Pdfcrowd;

class Example {
    public static void main(String[] args) throws Exception {
        Pdfcrowd.PdfToHtmlClient client =
            new Pdfcrowd.PdfToHtmlClient("demo", "demo");

        client.convertUrlToFile(
            "https://your-server.com/invoice.pdf", "invoice.html"
        );
    }
}

The URL must return a PDF and be reachable from PDFCrowd's servers. The client throws Pdfcrowd.Error on conversion or validation errors. See Handle Errors.

Convert a PDF File

Convert logo.pdf to HTML and save it locally as logo.html:

import com.pdfcrowd.Pdfcrowd;

class Example {
    public static void main(String[] args) throws Exception {
        Pdfcrowd.PdfToHtmlClient client =
            new Pdfcrowd.PdfToHtmlClient("demo", "demo");

        client.convertFileToFile("logo.pdf", "logo.html");
    }
}

Configure a Conversion

Authentication

Pass your PDFCrowd username and API key to PdfToHtmlClient. Find your credentials on the API Keys page.

Choose an Input

Send a PDF URL, upload a local file, or pass PDF bytes or a readable binary stream.

InputMethod
URLconvertUrlToFile()
Local fileconvertFileToFile()
Bytes in memoryconvertRawDataToFile()
Readable binary streamconvertStreamToFile()

Add Conversion Settings

Set conversion options on the client before calling a conversion method. For example, client.setScaleFactor(150); enlarges the HTML page layout by a factor of 1.5.

Common settings are listed below. See the method reference for all available options or browse Java examples.

PurposeMethods
Pages to convertsetPrintPageRange()
Password-protected inputsetPdfPassword()
Embedded or separate resourcessetImageMode(), setCssMode(), setFontMode()
Image format and resolutionsetImageFormat(), setDpi()
Scale and stylessetScaleFactor(), setCustomCss(), setHtmlNamespace()
Consistent ZIP outputsetForceZip()

The PDF password unlocks the input document; it is separate from your PDFCrowd API key.

Handle the Result

Choose an Output

By default, images, CSS and fonts are embedded in a single HTML file. The result preserves the PDF page layout; it does not reconstruct the original source HTML or create a responsive design.

The methods below use URL input; local files, bytes, and input streams have corresponding methods.

OutputMethod
Local fileconvertUrlToFile()
Byte arrayconvertUrl()
OutputStreamconvertUrlToStream()

File-output methods overwrite an existing destination file.

This example converts a PDF to HTML, receives the result as a byte[] and saves it as invoice.html:

import com.pdfcrowd.Pdfcrowd;
import java.nio.file.Files;
import java.nio.file.Path;

class Example {
    public static void main(String[] args) throws Exception {
        Pdfcrowd.PdfToHtmlClient client =
            new Pdfcrowd.PdfToHtmlClient("demo", "demo");

        byte[] result = client.convertUrl(
            "https://your-server.com/invoice.pdf"
        );
        Files.write(Path.of("invoice.html"), result);
    }
}

Setting an image, CSS or font mode to "separate" produces a ZIP archive with the HTML and its resources. Save that result with a .zip extension, then extract it while retaining the relative directory structure. isZippedOutput() reports whether the selected settings produce a ZIP archive.

When serving the result from a web application, use text/html; charset=utf-8 for HTML or application/zip for an archive.

Handle Errors

The client throws a Pdfcrowd.Error exception on conversion or validation errors. This example converts a PDF to HTML and logs any PDFCrowd error, including its HTTP status and reason code:

import com.pdfcrowd.Pdfcrowd;
import java.util.logging.Level;
import java.util.logging.Logger;

class Example {
    private static final Logger logger = Logger.getLogger(Example.class.getName());

    public static void main(String[] args) throws Exception {
        Pdfcrowd.PdfToHtmlClient client =
            new Pdfcrowd.PdfToHtmlClient("demo", "demo");

        try {
            client.convertUrlToFile(
                "https://your-server.com/invoice.pdf", "invoice.html"
            );
        } catch (Pdfcrowd.Error error) {
            logger.log(Level.SEVERE, "PDFCrowd conversion failed (status " +
                error.getStatusCode() + ", reason " + error.getReasonCode() + ")", error);
            throw error;
        }
    }
}

Local Java errors, such as a failure to read an input file or write the output, may need separate handling.

Pdfcrowd.Error provides these methods:

MethodReturns
getStatusCode()The HTTP status code, when available.
getReasonCode()The reason code identifying the specific error, or -1 if unavailable.
getMessage()The error message.
getDocumentationLink()A link to relevant documentation, when available.

Common Status Codes

StatusWhat it meansWhat to do
400Invalid input, settings, or conversion failureRead the reason code and correct the input or settings.
401Missing credentials or an inactive licenseCheck your username, API key, and license status.
403Suspended service or no credits remainingCheck your account and available credits.
413Upload exceeds the 300 MB limitReduce the upload size.
429Request rate limit reachedWait and reduce the rate of new requests.
430Concurrent request limit reachedAllow active requests to finish before starting more.
503Temporary network issueCheck your retry policy before submitting another request.

See all status and reason codes for specific explanations and Limits and Retries for retry behavior.

Read Conversion Information

This information is available after a conversion and describes the client's last conversion.

MethodUse
getJobId()Identify the conversion in logs and support requests.
getDebugLogUrl()The URL of the conversion debug log when logging is enabled.
getPageCount()Read the page count reported for the conversion.
getOutputSize()Read the HTML size in bytes.
getConsumedCreditCount()Read the credits consumed by the conversion.
getRemainingCreditCount()Read the remaining credit count reported with the conversion.

Limits and Retries

Request rate and concurrency limits depend on your license. Control how quickly your application submits conversions and how many it runs at once. A 429 response concerns request rate; a 430 response concerns requests already in progress. The maximum upload size is 300 MB.

The Java client automatically retries a request once when it receives HTTP 502 or 503. Use setRetryCount() to change that count, or set it to 0 to disable automatic retries. Account for these retries when adding an application-level retry policy.

Troubleshooting

Inspect a Conversion

This example converts a PDF to HTML with debug logging enabled and logs the debug log URL when available. Use the input and settings from the conversion you are investigating.

import com.pdfcrowd.Pdfcrowd;
import java.util.logging.Level;
import java.util.logging.Logger;

class Example {
    private static final Logger logger = Logger.getLogger(Example.class.getName());

    public static void main(String[] args) throws Exception {
        Pdfcrowd.PdfToHtmlClient client =
            new Pdfcrowd.PdfToHtmlClient("demo", "demo");

        try {
            client.setDebugLog(true);
            client.convertUrlToFile(
                "https://your-server.com/invoice.pdf", "invoice.html"
            );
        } catch (Pdfcrowd.Error error) {
            logger.log(Level.SEVERE, "PDFCrowd conversion failed (status " +
                error.getStatusCode() + ", reason " + error.getReasonCode() + ")", error);
            throw error;
        } finally {
            String debugLogUrl = client.getDebugLogUrl();
            if (debugLogUrl != null && !debugLogUrl.isEmpty()) {
                logger.info("Debug log: " + debugLogUrl);
            }
        }
    }
}

The debug log contains conversion settings and processing details. You can also find logs in your conversion history. A local or connection failure may occur before a conversion log is available.

Common Problems

ProblemCheck
The PDF cannot be loadedCheck that the URL returns a PDF, rather than HTML or a login page, and is reachable from PDFCrowd's servers. For local files, check the path and read permissions.
The PDF requires a passwordSet the input document's password with setPdfPassword().
Images, fonts or styles are missingIf resource modes are set to "separate", extract the ZIP archive and retain its relative directory structure.
The layout does not adapt to the screenThe HTML preserves the PDF page layout. Responsive layouts require further work on the generated HTML and CSS.

For help, contact support and include any available diagnostics, the client version, and enough detail to reproduce the problem.