Logo
Apdf tutorials October 2026 6 min read

Convert Word to PDF in Python — Without Microsoft Word on Your Server

The first answer you find for “python docx to pdf” is docx2pdf — which drives Microsoft Word, so it stops working the moment your code moves to a Linux server or a container. The usual workaround is installing a full office suite on the server: a large install, a subprocess to babysit, timeouts for the runs that hang, and fonts to manage.

If converting documents isn't your product, skip all of that. Send the Word file's URL to Apdf and get a PDF back — the same script runs on your laptop, a Linux server or a serverless function, with nothing installed besides requests.

What you'll build
A Python script that turns an agency's monthly client reports from Word into PDFs — one document at a time, then a whole batch without polling: report-acme-2026-09.docx → report-acme-2026-09.pdf  ·  1 page(s), 58670 bytes
Free plan: 500 PDF operations a month, no card
1

Convert one document

The endpoint takes the URL of the Word file and starts the conversion in the background, answering right away with a job_id. The script polls the job until the PDF is ready, then downloads it. Save it as convert_one.py:

import time
import requests

API_TOKEN = "YOUR_API_TOKEN"
BASE_URL = "https://apdf.io/api"
HEADERS = {"Authorization": f"Bearer {API_TOKEN}", "Accept": "application/json"}


def wait_for_job(job_id, timeout=300):
    """Poll the job status until the conversion has finished."""
    deadline = time.time() + timeout
    while time.time() < deadline:
        response = requests.post(f"{BASE_URL}/job/status/check", headers=HEADERS, data={"id": job_id})
        response.raise_for_status()
        job = response.json()
        if job["status"] == "successful":
            return job["result"]
        if job["status"] == "failed":
            raise RuntimeError(f"Conversion failed: {job}")
        time.sleep(2)
    raise TimeoutError(f"Job {job_id} did not finish in time")


def convert_to_pdf(document_url):
    """Start the conversion and return the finished PDF's details."""
    response = requests.post(f"{BASE_URL}/pdf/file/convert", headers=HEADERS, data={"file": document_url})
    response.raise_for_status()
    return wait_for_job(response.json()["job_id"])


if __name__ == "__main__":
    result = convert_to_pdf("https://files.northlight.example/reports/report-acme-2026-09.docx")
    print(f"{result['pages']} page(s), {result['size']} bytes: {result['file']}")

    pdf = requests.get(result["file"])
    pdf.raise_for_status()
    with open("report-acme-2026-09.pdf", "wb") as f:
        f.write(pdf.content)

The job status the script polls looks like this once the conversion is done — result is what wait_for_job returns:

{
    "id": "01M4D6SR4HQCWS53DGABFX9X4Z",
    "created": "2026-10-08T07:31:45.000000Z",
    "status": "successful",
    "result": {
        "file": "https://apdf-files.s3.eu-central-1.amazonaws.com/8914e0ea6ac746e16521e.pdf",
        "size": 59114,
        "pages": 1,
        "expiration": "2026-10-08T08:31:47.971497Z"
    }
}
Tip: The file has to be reachable by URL. If your documents sit in a private S3 bucket, pass a presigned URL (s3.generate_presigned_url in boto3) — it only needs to stay valid until the conversion has downloaded the file.

The PDF behind result["file"] is deleted after one hour (see expiration), so download it right away, as the script does, instead of storing the link.

2

Convert a whole batch with a webhook

Polling is fine for one file, but every status check is an API request — poll a dozen jobs side by side and you hit the rate limit. For a batch, let Apdf call you instead: pass a webhook_url with each document, and the finished job is POSTed to it. The query string tells the receiver which report the PDF belongs to.

from urllib.parse import quote

import requests

API_TOKEN = "YOUR_API_TOKEN"
BASE_URL = "https://apdf.io/api"
HEADERS = {"Authorization": f"Bearer {API_TOKEN}", "Accept": "application/json"}
WEBHOOK_URL = "https://reports.northlight.example/webhooks/converted"

documents = [
    "https://files.northlight.example/reports/report-acme-2026-09.docx",
    "https://files.northlight.example/reports/report-globex-2026-09.docx",
    "https://files.northlight.example/reports/report-initech-2026-09.docx",
]

for document_url in documents:
    name = document_url.rsplit("/", 1)[-1].rsplit(".", 1)[0]
    response = requests.post(
        f"{BASE_URL}/pdf/file/convert",
        headers=HEADERS,
        data={"file": document_url, "webhook_url": f"{WEBHOOK_URL}?name={quote(name)}"},
    )
    response.raise_for_status()
    print(f"Queued {name}: job {response.json()['job_id']}")

The receiver gets the same JSON as the job status — here is the one for the Acme report:

{
    "id": "01M4D6W4XV251DV2K796AAMT0G",
    "created": "2026-10-08T07:33:04.000000Z",
    "status": "successful",
    "result": {
        "file": "https://apdf-files.s3.eu-central-1.amazonaws.com/540ec6e56ac747300d8e4.pdf",
        "expiration": "2026-10-08T08:33:06.100458Z",
        "pages": 1,
        "size": 58670
    }
}

A minimal receiver with Python's standard library saves each PDF as it arrives. A failed job arrives with "status": "failed" and the error in result:

import json
from http.server import BaseHTTPRequestHandler, HTTPServer
from urllib.parse import parse_qs, urlparse

import requests


class ConvertedHandler(BaseHTTPRequestHandler):
    def do_POST(self):
        job = json.loads(self.rfile.read(int(self.headers["Content-Length"])))
        name = parse_qs(urlparse(self.path).query)["name"][0]

        if job["status"] == "successful":
            pdf = requests.get(job["result"]["file"])
            pdf.raise_for_status()
            with open(f"{name}.pdf", "wb") as f:
                f.write(pdf.content)
            print(f"✓ {name}.pdf ({job['result']['pages']} page(s))")
        else:
            print(f"✗ {name}: {job.get('result')}")

        self.send_response(200)
        self.end_headers()


HTTPServer(("0.0.0.0", 8000), ConvertedHandler).serve_forever()

Start the receiver, run the batch script, and the PDFs land in the receiver's folder within seconds:

Queued report-acme-2026-09: job 01M4D6W4XV251DV2K796AAMT0G
Queued report-globex-2026-09: job 01M4D6W4Z330YQ10QMPPDWEJX0
Queued report-initech-2026-09: job 01M4D6W50FEYQ3VD7XEDY0R7EF

✓ report-globex-2026-09.pdf (1 page(s))
✓ report-acme-2026-09.pdf (1 page(s))
✓ report-initech-2026-09.pdf (1 page(s))
Heads up: The webhook URL has to be reachable from the internet. While developing locally, expose the receiver with a tunnel such as ngrok or Cloudflare Tunnel and use that address as WEBHOOK_URL.
3

Reuse it for Excel and PowerPoint

Nothing in the scripts is specific to Word. The same endpoint converts Excel workbooks (.xlsx, .xls), PowerPoint decks (.pptx, .ppt), OpenDocument files (.odt, .ods, .odp), older .doc files and RTF — point documents at them and every page or slide becomes a PDF page.

Tip: Documents set in the standard Office fonts (Calibri, Cambria, Arial, Times New Roman) keep their line and page breaks. A custom brand font is replaced by a similar one, which can move text around — convert one sample before you automate a whole folder.

Where to go from here

The PDF is ready to send — and you can find out who actually reads it.

After the API call

Your code made the PDF.
Then it went dark.

Opened, read, re-read, dropped on page 4 — you never see any of it. Share the PDFs you generate through Apdf recipient links, and every signal becomes something you can act on: ping Slack, update the CRM, let an agent follow up. Same account, same API token, one more call.

Your PDF, after sending Live
document:loaded CFO
page:read p4 · 38s
link:clicked pricing

API · Webhook · MCP