Spreadsheets

Repair XLSX: Fix a Corrupted Excel File in Your Browser

Repair xlsx files in your browser: fix an Excel workbook that will not open, see what is damaged, and save the cells as CSV. Nothing is uploaded.

Use the Repair XLSX: Fix a Corrupted Excel File in Your Browser

Add an Excel file to see what is wrong with it

The page reads the file part by part, even when the ZIP package is cut off or damaged, repairs the XML inside it, fixes the links between parts, and writes a clean copy. It tells you what it found, what it did, and what could not be kept. If the main part cannot be repaired, it writes it again from what else is in the file.

Everything runs in your browser. What you enter is never uploaded or stored.

An .xlsx workbook looks like one file, but it is a ZIP package of many parts: one part for each sheet, one for the shared strings that hold the text of the cells, one for the cell styles, and lists that tell Excel what each part is and how the parts are linked. If a save was interrupted, a download was cut short, or a few bytes changed, one part or one link may be wrong, and Excel says that it found unreadable content, even though most of the numbers are still in the file.

Add the workbook on this page. It reads whatever can be read, repairs the XML in the sheets, fixes the links between the parts, and writes a clean copy. It tells you what was wrong, what it did, and what it could not keep, and it lists the cells it found, so that you can save any sheet as a CSV file even if the repaired workbook is not perfect. Nothing is uploaded, and the code behind the page makes no network requests, which matters for a workbook with figures that you are not meant to share.

How to use the Repair XLSX: Fix a Corrupted Excel File in Your Browser

  1. Add the workbookDrop the Excel file that will not open on the box, or choose it from your device. It is read on your device and does not leave it. Files up to 150 MB are accepted, in the .xlsx, .xlsm, and .xltx formats.
  2. Read what the page foundThe page says in one line whether the file was fine, repaired, or repaired with some loss. Below that it lists what was wrong, what was done, and what could not be kept, each with the part of the file it concerns, such as xl/worksheets/sheet2.xml.
  3. Check the cellsThe number of sheets, rows, filled cells, and formulas is shown, with the values the page found. If you only need the data, choose a sheet and download it as CSV. Dates show as the serial numbers that Excel stores, and formulas show the value they last had.
  4. Download the repaired workbookDownload the repaired .xlsx, open it in Excel, and compare it with the CSV. The page cannot test the file in Excel itself, so treat the first opening as the real check, and keep your original.

What is inside an .xlsx file

The Library of Congress describes an XLSX file as packaged with the Open Packaging Conventions, which are based on ZIP. Rename the file to .zip and you can open it. The top level of a minimal package has the folders _rels, docProps, and xl, and one file, [Content_Types].xml. The xl folder holds workbook.xml and a worksheets folder with one file for each sheet, plus the parts that support formatting, calculation order, and pictures.

Four kinds of part matter for repair. The sheets hold the cells as rows of cells with a value and a type. The shared strings part holds each different piece of text once, and a text cell holds only the number of its string, so a cell that points past the end of the list is an error. The styles part holds the cell formats, and a cell holds only the number of its format. The workbook part lists the sheets, each by a link ID, and the relationship files turn those IDs into part names. When any of these is missing or points to nothing, Excel reports unreadable content, even if every number is intact.

How a workbook gets damaged

The most common cause is a file that ends too soon: an interrupted download, an email attachment that was cut, a sync program that stopped halfway, a crash in the middle of a save. A ZIP file keeps its list of parts at the end, as the PKWARE format description explains, so a cut loses the list while the parts before the cut are still there. This page does not need the list. It finds each part by its header.

Other causes are changed bytes, which break the checksum of a compressed part, and programs that write a file badly. The first leaves a sheet with some rows wrong or missing. The second leaves a tag unclosed, an ampersand unescaped, or a cell that points at a style or a string that is not in the file. Each of these has its own repair below. One more detail is worth knowing: some programs write the workbook part near the end of the file, and a cut then takes the workbook part, and with it the names of the sheets. The page writes the workbook again from the sheets that are left, and names them Sheet1, Sheet2, and so on.

How the file is read

Every part of a ZIP file is compressed with the DEFLATE method, described in RFC 1951. The browser's own decompressor throws away the output of a stream that turns out to be damaged. This page has its own decoder, which keeps everything it decoded up to the point of the damage and says where it stopped. A sheet that was cut off in the middle is therefore not lost: its first rows are used.

Each part is checked against the size and checksum that the file recorded. A part that matches is in order. A part that decodes but does not match is marked as changed. A part that stops early is marked as cut off. In a sheet that is cut off, the row that was being written is dropped, because half a row would show wrong data, such as a name with its amount missing. Where compressed data is damaged, the last row before the point where decoding stopped is dropped too, because the bytes just before the damage are often wrong.

What is repaired

XML parts are read by a tolerant reader and written again as strict XML. It closes elements that were left open at a cut, removes end tags that close nothing, repairs an end tag whose name was damaged by a character or two, escapes ampersands that are not part of an entity, removes characters that XML does not allow, and declares again the namespace prefixes that a damaged part lost.

Excel has its own rules on top of XML, and the page checks the ones that a cut or a lost part breaks. A text cell that points at a shared string that is not in the file is emptied. A cell whose number is not a number is emptied. A cell, row, or column that names a style that is not in the styles part, or that has no styles part to point at, loses its format and keeps its value. A count of merged cells or of validations that does not match the entries is corrected. A link to a sheet, picture, or drawing that is not in the file is removed, with the sheet entry or drawing that used it. If a sheet is removed, the defined names go, because they refer to sheets by position.

The formula order list, xl/calcChain.xml, is left out whenever the workbook is repaired. Excel builds it again when it opens the file, and it reports a problem if the old list names cells that are no longer in the file. The formulas themselves are kept.

When the workbook part cannot be repaired

If the workbook part cannot be read as XML, or the list of its links is gone, the page writes both again from the sheets it found, with their styles, shared strings, and theme linked. When the old workbook part could still be read, the sheet names are kept, and otherwise they are numbered. What the workbook as a whole holds, such as defined names and calculation settings, is not in the new one, and the page says so. If no sheet can be read, the page says that nothing could be recovered.

A file that is not an .xlsx is told apart and not guessed at. A Word or PowerPoint file, an older .xls file, a password-protected file, and a PDF each get their own message. The older .xls format is a different kind of file, not a ZIP package, and this page does not repair it. In Excel, try File, Open, choose the file, press the arrow next to Open, and choose Open and Repair.

How the repair was tested

Three workbooks were made with libraries that are not part of this page. One, made with openpyxl, has three sheets, formulas, merged cells, a defined name, a table, a picture, a comment, special characters, and a data sheet of 3,000 rows. A copy of it has a formula order list added, which Excel writes and openpyxl does not. The third was made with XlsxWriter, which stores shared strings and the results of formulas and puts the workbook part early in the file as Excel does. The workbooks were each tried undamaged and then damaged in 59 ways in all: cut inside a sheet at three points, with the list of parts removed, with the ending cut off, with the workbook part wrecked or its links missing, with sheets, styles, or pictures removed, with the shared strings cut off or gone, with four kinds of XML damage in a sheet, and with single bytes flipped in compressed data.

The 57 repaired workbooks were opened with openpyxl, tested with Python's zipfile, and parsed with lxml's strict XML reader. All passed. The cells that openpyxl reads were the real beginning of the original cells, and the numbers of sheets, rows, and cells that the page listed matched what openpyxl read. The exceptions are the cases where content was lost on purpose or by damage: after flipped bytes, the text of some rows is different, and after the shared strings are cut, cells whose text is gone are empty. A last test damaged the samples in random ways, 60 times, and every file was either repaired and passed its own checks, or refused with a clear message.

Limits and accuracy

  • The repaired workbook was not opened in Microsoft Excel, because Excel was not available for the tests. It was read by openpyxl, lxml, and zipfile. Open it in Excel, and keep your original.
  • The older .xls format is not supported. It is a different kind of file, and this page repairs .xlsx workbooks, which are ZIP packages. Files protected with a password are not repaired either.
  • Where compressed data is damaged in the middle, the rows after the damage are lost, and the text of a few rows before it may be wrong. A checksum that does not match is reported, but the page cannot tell which cell is wrong.
  • A cell whose text, number, or format is gone is left empty or unformatted. The page lists these as repairs, but it cannot bring the content back.
  • Pictures and drawings that are cut off are removed. Pivot tables, charts, and external links are kept only if their parts are in the file, and the page does not check their rules.
  • The page checks the package, the XML, the links, and the rules for cells that are listed above. It does not check every rule of the spreadsheet schema, so a file that is sound by these checks may still be refused by Excel.
  • The CSV and the text show dates as the serial numbers that Excel stores and formulas as the value they last had. A formula that was never worked out shows an empty cell. Sheets are listed up to 200,000 rows each.
  • A macro-enabled file (.xlsm) is repaired as a package, and its macros are kept as they are. The page does not run or check them.

Frequently asked questions

How do I repair a corrupted XLSX file?

Add the workbook on this page. It reads the parts of the package, repairs the XML in the sheets and the links between the parts, and writes a clean copy that you download. It lists what was wrong and what could not be kept, and you can save any sheet as CSV.

Why does Excel say that it found unreadable content?

Excel refuses the whole workbook if one list is missing, one link points to nothing, one cell points at a style or string that is not there, or one part is not well-formed XML, even when the numbers are all intact. Fixing the structure often brings the workbook back.

Can it recover a workbook that was cut off during a download or a save?

Often, yes. The ZIP list of parts is at the end of the file, so a cut loses it, but the parts before the cut can still be read. A sheet cut in the middle gives its rows up to the cut, without the half-written last row.

Will my formulas and formatting be kept?

Formulas and formats in the parts that were recovered stay. A cell that points at a style or string that is gone loses that style or text and keeps its value. The page does not rebuild the formula order list, because Excel does that when the file opens.

Does it work on .xls files?

No. An .xls file is the older binary format, not a ZIP package. This page repairs .xlsx workbooks. For an .xls file, try Excel's Open and Repair command: choose File, Open, select the file, press the arrow next to Open, and choose Open and Repair.

Can I get my data out if the repair is not perfect?

Yes. The page lists the cells it found even when a repaired workbook is not offered, and you can download any sheet as a CSV file. Dates appear as the serial numbers that Excel stores, so format that column as a date in the program you open the CSV in.

Is my workbook uploaded to a server?

No. The file is read and repaired in your browser, and the code behind the page makes no network requests. The repaired copy and the CSV are created on your device and downloaded from there.

Research and references

This page was written and checked against the sources below.

  1. Library of Congress: XLSX Transitional (Office Open XML), ISO 29500, ECMA-376
  2. Ecma International: ECMA-376, Office Open XML file formats
  3. PKWARE: ZIP file format specification (APPNOTE)
  4. RFC 1951: DEFLATE Compressed Data Format Specification
  5. Microsoft Learn: Spreadsheets with the Open XML SDK