A .docx file looks like one document, but it is a ZIP package of many small parts: the text, the styles, the pictures, the headers, and the lists that tell Word what each part is and how the parts are linked. If a save was interrupted, a download was cut short, or a few bytes changed, one of those parts or links may be wrong, and Word says that the file is corrupt, even though most of the text is still in it.
Add the file on this page. It reads whatever can be read, repairs the XML inside the parts, fixes the links between them, and writes a clean copy. It tells you what was wrong, what it did, and what it could not keep, and it shows the text it found. If the document cannot be repaired, it rescues the text. Nothing is uploaded, and the code behind the page makes no network requests, which matters for a document you are not meant to share.
How to use the DOCX Repair: Fix a Corrupted Word File in Your Browser
- Add the fileDrop the Word file that will not open on the box, or choose it from your device. It is read on your device and does not leave it. Files up to 150 MB are accepted.
- Read what the page foundThe page says in one line whether the file was fine, repaired, repaired with some loss, or rebuilt from its text. Below that it lists what was wrong, what was done, and what could not be kept, each with the part of the file it concerns.
- Check the textThe text found in the file is shown, with the number of paragraphs, words, tables, and pictures. Compare it with what you remember. If the text is what you need, you can download it as a .txt file even if the repaired document is not perfect.
- Download the repaired documentDownload the repaired .docx. Open it in Word, check it, and keep your original. The page cannot test the file in Word itself, so treat the first opening as the real check.
What is inside a .docx file
The Library of Congress describes a DOCX file as packaged with the Open Packaging Conventions, which are based on ZIP. Rename a .docx to .zip and you can open it. The top level of a minimal package has the folders _rels, docProps, and word, and one file, [Content_Types].xml. The word folder holds the main content in document.xml. Headers and footers are separate parts. Pictures sit in a media folder.
Two small files hold the package together. [Content_Types].xml is mandatory in any package and lists the content type of the parts, so that Word knows a part is a style sheet and not a picture. The relationship files, called .rels, list the links between parts. The one in the _rels folder points to word/document.xml as the main document, and the one beside the document points to its styles, settings, headers, and pictures, each by an ID that the text of the document refers to. When any of this is missing or points to nothing, Word refuses the file, even if every paragraph is intact.
How a Word file gets damaged
The most common cause is a file that ends too soon: an interrupted download, an email attachment cut by a mail gateway, a sync program that stopped halfway, a crash in the middle of a save. A ZIP file keeps its list of parts at the end, as the PKWARE format description explains, so a file that is cut off loses the list while the parts before the cut are still there. Most programs give up. This page does not need the list. It finds each part by its header, which comes in front of the part's data.
The other causes are changed bytes and bad structure. A byte that was changed in a compressed part breaks the checksum, and may break the compressed data from that point. A program that wrote the file badly may leave a tag unclosed, an ampersand unescaped, or a table cell with no paragraph, which Word does not accept. A file may arrive with extra bytes in front, such as a mail header. Each of these has its own repair below.
How the file is read
Every part in a ZIP file is compressed with the DEFLATE method, described in RFC 1951. The browser's own decompressor throws away the output of a stream that turns out to be damaged. This page has its own decoder, which keeps everything it decoded up to the point of the damage and says where it stopped. A part that was cut off in the middle is therefore not lost: its first half is used.
Each part is checked against the size and checksum that the file recorded. A part that matches is marked as in order. A part that decodes but does not match is marked as changed. A part that stops early is marked as cut off. Where the compressed data itself is damaged, the last paragraph before the point where decoding stopped is dropped, because the bytes just before the damage are often wrong too. A part that was cut off cleanly keeps its last, shorter paragraph.
What is repaired
XML parts are read by a tolerant reader and written again as strict XML. It closes elements that were left open at a cut, removes end tags that close nothing, and repairs an end tag whose name was damaged by a character or two. It escapes ampersands that are not part of an entity, removes characters that XML does not allow, puts quotes around attribute values that have none, and drops a tag that was cut off at the end. It declares again the namespace prefixes that a damaged part lost, using the addresses that Office files use, and it reads text saved as UTF-16.
Word has its own rules on top of XML. A table cell must end with a paragraph, a row must have cells, and a table must have rows. A save that was cut off in the middle of a table breaks all three, so the page adds the empty paragraphs and removes the empty rows. A link to a part that is not in the file is removed, and so is the picture, header, or footer that used it, because Word refuses a document that refers to a link that does not exist. Missing content types and package relationships are written again from the parts that are there.
When the document cannot be repaired
If the main part cannot be read as a Word document at all, the page looks through every part for text and puts it, paragraph by paragraph, into a new document with no formatting. This is the last resort, and the page says so with a different headline: the formatting, tables, and pictures of the original are not in it. The text is also offered as a .txt file. If no text can be found anywhere, the page says so, and the file is not changed.
A file that is not a .docx is told apart and not guessed at. An old .doc file, a password-protected file, an RTF file with a .docx name, a PDF, and an Excel or PowerPoint file each get their own message. For an old .doc file, the message points to Word's own Open and Repair command.
How the repair was tested
Four Word files were made with python-docx: a report with headings, bold and italic text, special characters, a table, a bulleted list, a header, a footer, and a picture; a long document of 300 paragraphs; and two copies of the report, one written as a stream with data descriptors and one with every part stored uncompressed. They were damaged in 55 ways: cut inside the document at five points, with the list of parts removed, with the ending cut off, with parts or links removed, with ten kinds of XML damage, with single bytes flipped in the compressed text, and with a mail header in front.
The 47 repaired files were opened with python-docx, tested with Python's zipfile, and parsed with lxml's strict XML reader, none of which is part of this page. All 47 passed. The paragraphs they held were the real beginning of the original, and the more of a cut-off file there was, the more came back. In the cut files the number of paragraphs recovered rose from 3 to 17 of 18 as the cut moved toward the end of the document. In the four cases where the cut fell inside the first tag of the document, before any text, the page said that nothing could be recovered. The decoder was compared with zlib on random and text data at every compression level, and a last test damaged the samples in random ways, 60 times, and every file was either repaired and passed its own checks, or refused with a clear message.
Limits and accuracy
- The repaired file was not opened in Microsoft Word, because Word was not available for the tests. It was read by python-docx, lxml, and zipfile. Open it in Word, and keep your original.
- Where compressed data is damaged in the middle, the text after the damage is lost, because the compression cannot be followed past it. Only the text before the damage is recovered.
- A picture or other part that is cut off is left out, together with the place in the text that showed it. A part whose checksum is wrong is kept, and may look wrong.
- The page checks the package, the XML, the links, and the table rules. It does not check every rule of the Word schema, so a file that is sound by these checks may still be refused by Word.
- Old Word files with the .doc ending, and files protected with a password, are not repaired. The page says which one it was given.
- When the document is rebuilt from its text, the styles, tables, pictures, headers, and footers of the original are not in the new document.
- Files larger than 4 GB, which use the ZIP64 extension, are not supported, and the page works on files up to 150 MB.
- A macro-enabled file (.docm) is repaired as a package, and its macros are kept as they are. The page does not run or check them.
Frequently asked questions
How do I repair a corrupted DOCX file?
Add the file on this page. It reads the parts of the package, repairs the XML and the links between the parts, and writes a clean copy that you download. It lists what was wrong and what could not be kept. If the document cannot be repaired, it gives you the text.
Why does Word say that my file is corrupt when most of the text is still there?
A .docx file is a package of parts and lists that link them. Word refuses the whole file if one list is missing, one link points to nothing, or one part is not well-formed XML, even when the text is complete. Fixing the structure often brings the document back.
Can it recover a file that was cut off during download or save?
Yes, often. A ZIP file keeps its list of parts at the end, so a cut loses the list and everything after the cut. The parts before the cut can still be read, and the page reads them by their headers. A part that was cut in the middle gives the text up to the cut.
Will the formatting and pictures be kept?
When the document part is repaired, yes: the styles, tables, headers, and pictures that were recovered stay. A picture that was cut off is removed. If the document had to be rebuilt from its text, only the text is kept, and the page says so in the headline.
Does it work on old .doc files?
No. The old .doc format is a different kind of file. This page repairs the newer .docx package. For a .doc file, try Word's Open and Repair command: choose File, Open, select the file, press the arrow next to Open, and choose Open and Repair.
Can I get the text if the repair fails?
If any text can be found in the file, the page shows it and lets you download it as a .txt file, even when the repaired document is not offered. If no text can be found in any part, the page says that nothing could be recovered.
Is my document uploaded to a server?
No. The file is read and repaired in your browser, and the code behind the page makes no network requests. The repaired copy is created on your device and downloaded from there.
Research and references
This page was written and checked against the sources below.

