File formats that are secretly ZIP files
Most people meet ZIP as the .zip file on their desktop, but the format's real victory is quieter: it became the container that other formats are built on. The contract you signed in Word (.docx), the app you installed on your phone (.apk), the ebook you are reading (.epub), the Java library your project depends on (.jar), the Python wheel your build just downloaded (.whl) — every one of them is a ZIP archive underneath, with an agreed internal layout riding on top. The first two bytes of all of them are the letters PK, the initials of the format's inventor, Phil Katz.
This page is a field guide to that family: which formats are ZIP in disguise, what each one stores inside the container, the specification details that make them behave the way they do, and how to inspect any of them in your browser with no upload. It complements the broader archive formats overview and borrows its vocabulary from the glossary.
Why so many formats are built on ZIP
A file format needs a container long before it needs compression: named parts, folder hierarchy, a way to locate one part without reading the rest, and a checksum per entry. Building that from scratch is a year of engineering and a decade of security patches. ZIP already had all of it, on every operating system, by the mid-1990s — so format designers inherited the container instead of reinventing it.
Three container features do most of the recruiting work. Per-entry compression choice means a format can store one small file uncompressed — because a reader must identify the file before inflating anything — and deflate the rest (the store method vs Deflate split). A central directory at the end of the file gives random access to any part. And a CRC32 per entry means every part carries its own integrity check.
- Named entries with folder hierarchy encoded in entry paths — a package layout for free.
- Per-entry compression: each file can be stored or deflated independently.
- A central directory enables random access; a CRC32 per entry verifies each part.
- A ZIP reader ships with every OS since the 1990s — formats built on it inherit the tooling.
Office documents: docx, xlsx, pptx
Since Office 2007, Word, Excel, and PowerPoint save the Office Open XML format (ECMA-376, later ISO/IEC 29500) — a ZIP whose parts are XML plus embedded media. The skeleton is always the same: [Content_Types].xml declares the content type of every part, _rels/.rels wires up the relationships (including which part is the main document), and word/document.xml — or ppt/slides/slide1.xml, xl/workbook.xml — carries the actual body.
You can prove it in ten seconds: copy a .docx, rename the copy to .zip, and open it. What you see is the honest anatomy of the format — styles.xml, settings.xml, a media/ folder for images — and also the limit of the rename trick: an XML soup you can browse but not comfortably read. A viewer that understands the package goes one layer deeper: ZIPTool opens a .docx as flowing text, an .xlsx as a grid, a .pptx as slides, rather than a file listing — see the docx viewer page.
- docx/xlsx/pptx have been ZIPs of XML parts since Office 2007 (OOXML, ECMA-376 / ISO/IEC 29500).
[Content_Types].xmland the_rels/hierarchy define the package;word/document.xmlis the body.- Renaming a copy to .zip shows the parts — a package-aware viewer shows the document.
OpenDocument: odt, ods, odp
The open alternative to OOXML — OpenDocument, standardized by OASIS as ISO/IEC 26300 and used by LibreOffice — is also a ZIP, with one detail that became the family's most-quoted rule: the very first entry must be a file named mimetype, stored with no compression, so a reader can identify the document type from a few fixed bytes without inflating anything. After it come META-INF/manifest.xml (the part list) and content.xml (the body).
In an archive tool, an .odt opens as the container listing: content.xml, styles.xml, META-INF/ — all previewable as text. ZIPTool lists the container and previews those XML parts; it does not render ODF as a paginated document, which is the honest limit of the archive view.
- odt/ods/odp follow ODF (ISO/IEC 26300): a ZIP of
content.xml,styles.xml, and aMETA-INF/manifest. - Rule: the first entry is
mimetype, stored uncompressed — identifiable before any inflation. - Archive tools show the parts; document-level rendering is a separate capability.
Apps and ebooks: APK, EPUB, ipa, xpi, cbz
An EPUB ebook is a ZIP that borrows ODF's mimetype rule verbatim — first entry, stored, uncompressed — then adds its own packaging: META-INF/container.xml names the .opf package file, which defines the manifest of documents and images and the reading order (the spine). More on the format in what is an EPUB file.
An Android .apk is a ZIP laid out as an app: AndroidManifest.xml (name, permissions, entry points — encoded as compact binary XML, not the text XML you might expect), classes.dex (the bytecode Android executes), resources.arsc (the compiled resource table), and per-CPU-architecture lib/ folders of native libraries. Android also aligns entries on 4-byte boundaries so the OS can memory-map them cheaply. More in what is an APK file.
The shape repeats across ecosystems: an iOS .ipa is a ZIP whose first folder is Payload/ containing the .app bundle; a Firefox extension .xpi is a ZIP with an install manifest; a .cbz comic is literally a renamed ZIP of page images.
- EPUB: the mimetype-first rule inherited from ODF;
META-INF/container.xmlpoints at the OPF spine. - APK: binary-encoded manifest,
classes.dexcode,resources.arsc, per-ABI native libraries, zip-aligned entries. - ipa (
Payload/*.app), xpi (browser extensions), cbz (comics) — the same container across ecosystems.
Developer packaging: JAR, wheels, and cousins
A Java .jar is a ZIP with a contract: META-INF/MANIFEST.MF carries the metadata — including Main-Class, the entry point that makes a jar runnable — and a signed jar adds .SF and .RSA signature files under the same META-INF/ folder. Because class files are ordinary entries, a jar doubles as a classpath, and decompiling one is a matter of reading entries and running the bytecode through a decompiler, which ZIPTool does entirely in the browser (Java decompiler).
A Python .whl is a ZIP whose filename is load-bearing — it encodes the package name, version, and compatibility tags — and whose contents include a dist-info/ folder with METADATA, WHEEL, and RECORD, where RECORD lists every file in the wheel together with its hash: an integrity manifest riding inside the container. See what is a WHL file. The pattern continues wherever developers move named bundles: NuGet .nupkg, Java EE .war/.ear, AR scenes in .usdz, 3D-printing .3mf.
- jar:
META-INF/MANIFEST.MF(withMain-Class) makes a ZIP runnable as a Java program. - whl: the filename encodes install tags;
dist-info/RECORDhashes every file inside. - nupkg, war/ear, usdz, 3mf — one container serving many ecosystems.
How to tell whether a file is secretly a ZIP
The signature is unambiguous when the file starts with the ZIP structure: the two bytes 50 4B (PK, for Phil Katz) followed by 03 04 — the local file header magic. An empty archive instead opens with 50 4B 05 06 (the end-of-central-directory record), and a spanned archive with 50 4B 07 08. Practically every format on this page starts at offset 0 with 50 4B 03 04.
One exception is worth knowing: self-extracting installers prepend a program stub and append the ZIP after it, so the PK signature sits in the middle of the file rather than at the start. And none of the checks need special tools. On macOS or Linux, unzip -l report.docx lists the entries — if a tool built for zips lists the file's parts, it is a zip. On Windows, rename a copy to .zip and open it. Or skip the rename entirely and drop the file into a viewer that accepts the whole family: ZIPTool opens .apk, .epub, .jar, .whl, .ipa, .odt and the rest directly and lists their entries with no upload (why that matters).
50 4B 03 04at offset 0 = a ZIP local file header;50 4B 05 06= an empty archive.- Self-extracting archives hide the signature mid-file, behind the stub program.
- Check with
unzip -l, a renamed copy — or by dropping the file into an in-browser viewer.
The rename trick, and where it stops being useful
Renaming .docx to .zip is the classic move — the IT-support answer to “this file looks corrupt, let me look inside”. It works because the extension is only a hint; the bytes are the format. But the trick stops at the container layer: you get word/document.xml, not the document. Reading the actual contract still needs a package-aware viewer, and for Office files ZIPTool makes that the default — drop a .docx and it renders as text; rename the same bytes to .zip and it lists the package. Same file, two honest views.
Repacking is where care is needed. Zipping a folder back into an .epub or .odt with a generic archiver can produce a file strict readers reject, because the mimetype entry must stay first and uncompressed — a detail most archivers' defaults break silently. Office packages are more forgiving about entry order but not about missing parts: [Content_Types].xml and the _rels/ hierarchy are load-bearing. Edit on a copy, and check the result in the target application before deleting the original.
- The extension is a hint; the bytes are the format — renaming changes the hint, not the file.
- The rename shows the package, not the document; package-aware viewers go one layer deeper.
- Repack carefully: mimetype-first (EPUB/ODF) and
[Content_Types].xml(Office) are load-bearing.
Frequently asked questions
Is a .docx file a ZIP file?
Yes. A .docx created by Word 2007 or later is a ZIP archive containing XML parts — [Content_Types].xml, relationship files, and word/document.xml for the body text. Copy the file, rename the copy to .zip, and any zip tool opens it. The DOC name survives for compatibility; the container underneath is ZIP.
Which file formats are secretly ZIP archives?
The common ones: Office OOXML documents (docx, xlsx, pptx and their templates), OpenDocument files (odt, ods, odp), EPUB ebooks, Android app packages (apk), iOS apps (ipa), Firefox extensions (xpi), Java archives (jar, war, ear), Python wheels (whl), comic archives (cbz), NuGet packages (nupkg), and scene or printing formats like usdz and 3mf. They all start with the bytes 50 4B 03 04.
Can I open an APK or EPUB with a zip extractor?
Yes — both are ZIP containers, and ZIPTool accepts them directly with no renaming: an .epub lists its chapters and images, an .apk lists its manifest, code, and resources. One expectation to set: some parts, like AndroidManifest.xml inside an APK, are stored in a compact binary encoding rather than readable text XML, so those entries preview as binary.
Why is the mimetype file in an EPUB or ODT stored uncompressed?
The ODF and EPUB specifications require mimetype to be the first entry in the archive and stored with no compression, so a reader can identify the document type by reading a few fixed bytes — without inflating anything or scanning the whole file. It trades a handful of uncompressed bytes for instant format detection, and it is exactly the detail most generic re-zipping tools break.
Is it safe to edit or re-zip one of these files?
With care, and on a copy. The container is an ordinary ZIP, so adding or removing entries works — but the internal layout is a contract: EPUB and ODT need mimetype first and uncompressed, Office packages need [Content_Types].xml and their _rels files intact, and a wheel's RECORD hashes must match its files. Keep the original until the edited file opens correctly in its target application.
How can I check whether an unknown file is a ZIP?
Look at the first bytes: a ZIP archive begins with 50 4B 03 04 — the letters PK followed by 03 04. On macOS or Linux, run unzip -l on the file: if it lists entries, the file is a ZIP regardless of its extension. The exception is self-extracting installers, where a program stub precedes the archive and the PK signature sits mid-file. Dropping the file into an in-browser archive viewer is the zero-install check.