Convert Word to PDF in Python — Without Microsoft Word on Your Server
The first answer you find for “python docx to pdf” is docx2pdf
— which drives Microsoft Word, so it stops working the moment your code moves to a Linux
server or a container. The usual workaround is installing a full office suite on the server: a large install,
a subprocess to babysit, timeouts for the runs that hang, and fonts to manage.
If converting documents isn't your product, skip all of that. Send the Word file's URL to
Apdf and get a PDF back — the same script runs on your laptop, a Linux server
or a serverless function, with nothing installed besides requests.
Convert one document
The endpoint takes the URL of the Word file and starts the conversion in the background,
answering right away with a job_id. The script polls
the job until the PDF is ready, then downloads it. Save it as
convert_one.py:
import time
import requests
API_TOKEN = "YOUR_API_TOKEN"
BASE_URL = "https://apdf.io/api"
HEADERS = {"Authorization": f"Bearer {API_TOKEN}", "Accept": "application/json"}
def wait_for_job(job_id, timeout=300):
"""Poll the job status until the conversion has finished."""
deadline = time.time() + timeout
while time.time() < deadline:
response = requests.post(f"{BASE_URL}/job/status/check", headers=HEADERS, data={"id": job_id})
response.raise_for_status()
job = response.json()
if job["status"] == "successful":
return job["result"]
if job["status"] == "failed":
raise RuntimeError(f"Conversion failed: {job}")
time.sleep(2)
raise TimeoutError(f"Job {job_id} did not finish in time")
def convert_to_pdf(document_url):
"""Start the conversion and return the finished PDF's details."""
response = requests.post(f"{BASE_URL}/pdf/file/convert", headers=HEADERS, data={"file": document_url})
response.raise_for_status()
return wait_for_job(response.json()["job_id"])
if __name__ == "__main__":
result = convert_to_pdf("https://files.northlight.example/reports/report-acme-2026-09.docx")
print(f"{result['pages']} page(s), {result['size']} bytes: {result['file']}")
pdf = requests.get(result["file"])
pdf.raise_for_status()
with open("report-acme-2026-09.pdf", "wb") as f:
f.write(pdf.content)
The job status the script polls looks like this once the conversion is done —
result is what wait_for_job returns:
{
"id": "01M4D6SR4HQCWS53DGABFX9X4Z",
"created": "2026-10-08T07:31:45.000000Z",
"status": "successful",
"result": {
"file": "https://apdf-files.s3.eu-central-1.amazonaws.com/8914e0ea6ac746e16521e.pdf",
"size": 59114,
"pages": 1,
"expiration": "2026-10-08T08:31:47.971497Z"
}
}
s3.generate_presigned_url in boto3) —
it only needs to stay valid until the conversion has downloaded the file.
The PDF behind result["file"] is deleted after one hour
(see expiration), so download it right away, as the script does,
instead of storing the link.
Convert a whole batch with a webhook
Polling is fine for one file, but every status check is an API request — poll a dozen
jobs side by side and you hit the rate limit.
For a batch, let Apdf call you instead: pass a webhook_url
with each document, and the finished job is POSTed to it. The query string tells the receiver
which report the PDF belongs to.
from urllib.parse import quote
import requests
API_TOKEN = "YOUR_API_TOKEN"
BASE_URL = "https://apdf.io/api"
HEADERS = {"Authorization": f"Bearer {API_TOKEN}", "Accept": "application/json"}
WEBHOOK_URL = "https://reports.northlight.example/webhooks/converted"
documents = [
"https://files.northlight.example/reports/report-acme-2026-09.docx",
"https://files.northlight.example/reports/report-globex-2026-09.docx",
"https://files.northlight.example/reports/report-initech-2026-09.docx",
]
for document_url in documents:
name = document_url.rsplit("/", 1)[-1].rsplit(".", 1)[0]
response = requests.post(
f"{BASE_URL}/pdf/file/convert",
headers=HEADERS,
data={"file": document_url, "webhook_url": f"{WEBHOOK_URL}?name={quote(name)}"},
)
response.raise_for_status()
print(f"Queued {name}: job {response.json()['job_id']}")
The receiver gets the same JSON as the job status — here is the one for the Acme report:
{
"id": "01M4D6W4XV251DV2K796AAMT0G",
"created": "2026-10-08T07:33:04.000000Z",
"status": "successful",
"result": {
"file": "https://apdf-files.s3.eu-central-1.amazonaws.com/540ec6e56ac747300d8e4.pdf",
"expiration": "2026-10-08T08:33:06.100458Z",
"pages": 1,
"size": 58670
}
}
A minimal receiver with Python's standard library saves each PDF as it arrives. A failed job
arrives with "status": "failed" and the error in
result:
import json
from http.server import BaseHTTPRequestHandler, HTTPServer
from urllib.parse import parse_qs, urlparse
import requests
class ConvertedHandler(BaseHTTPRequestHandler):
def do_POST(self):
job = json.loads(self.rfile.read(int(self.headers["Content-Length"])))
name = parse_qs(urlparse(self.path).query)["name"][0]
if job["status"] == "successful":
pdf = requests.get(job["result"]["file"])
pdf.raise_for_status()
with open(f"{name}.pdf", "wb") as f:
f.write(pdf.content)
print(f"✓ {name}.pdf ({job['result']['pages']} page(s))")
else:
print(f"✗ {name}: {job.get('result')}")
self.send_response(200)
self.end_headers()
HTTPServer(("0.0.0.0", 8000), ConvertedHandler).serve_forever()
Start the receiver, run the batch script, and the PDFs land in the receiver's folder within seconds:
Queued report-acme-2026-09: job 01M4D6W4XV251DV2K796AAMT0G
Queued report-globex-2026-09: job 01M4D6W4Z330YQ10QMPPDWEJX0
Queued report-initech-2026-09: job 01M4D6W50FEYQ3VD7XEDY0R7EF
✓ report-globex-2026-09.pdf (1 page(s))
✓ report-acme-2026-09.pdf (1 page(s))
✓ report-initech-2026-09.pdf (1 page(s))
WEBHOOK_URL.
Reuse it for Excel and PowerPoint
Nothing in the scripts is specific to Word. The same endpoint converts Excel workbooks
(.xlsx, .xls),
PowerPoint decks (.pptx, .ppt),
OpenDocument files (.odt, .ods,
.odp), older .doc files and RTF —
point documents at them and every page or slide becomes a PDF page.
Where to go from here
The PDF is ready to send — and you can find out who actually reads it.
Related tutorials
Add Page Numbers and Headers to PDFs with Python
Add consistent page numbers, headers, or footers to existing PDF documents using Python and the Apdf Overlay API. Perfect for preparing reports for distribution without regenerating the original documents.
Convert Scanned PDFs to Searchable Documents with Python
Transform image-based scanned PDFs into fully searchable documents using Python and the Apdf OCR API. Perfect for digitizing paper archives, enabling text selection, and making legacy documents ready for modern search systems.
Fix Sideways PDF Scans: Auto-Rotate Pages with Python
Learn how to automatically fix sideways or upside-down PDF pages from mobile uploads. Using Python and the Apdf Rotate API, you can detect and correct orientation issues in scanned documents, making them ready for review workflows.
Your code made the PDF.
Then it went dark.
Opened, read, re-read, dropped on page 4 — you never see any of it. Share the PDFs you generate through Apdf recipient links, and every signal becomes something you can act on: ping Slack, update the CRM, let an agent follow up. Same account, same API token, one more call.
API · Webhook · MCP