A file name is a label that anyone can change. Renaming a PNG to .jpg does not make it a JPEG, and renaming a program to .pdf does not make it a document. Programs that open files look at the first bytes instead, and so does this page. It tells you what a file really is, whether its name agrees, and whether anything unusual comes after the end of the file.
This is useful when a download will not open, when a file has no extension, when a web page gave you something other than what you expected, or when an attachment looks wrong. It is also a quick check on a file you were sent. Everything runs in your browser. The page reads only the start and the end of each file, nothing is uploaded, and the code behind the page makes no network requests.
How to use the Check File Type Online: Find What a File Really Is
- Add the filesDrop one or more files on the box, or choose them from your device. Any type works, and up to 200 files can be checked at once. A large file is read only at its start and its end.
- Read the answerEach file shows its real type and MIME type, and whether the name agrees with the contents. A file that matches is marked green. A note means the name is more general or less exact. A mismatch or an unsafe look is marked in yellow or red.
- Look at the findingsFindings below the answer say what else was found: a program or archive attached to a picture, a file that looks cut off, macros in an Office file, or a name that hides its real extension.
- See how it was decidedOpen the details to see the bytes or the names inside the file that led to the answer, how certain it is, the usual extensions, and the first 128 bytes in hex.
- Rename or exportCopy the suggested name and rename the file yourself. For many files, show only those that need a second look, or download the results as a CSV file.
Why the name is not enough
The extension is part of the name, and the name can say anything. Operating systems use it to choose a program to open the file, and that makes it a convenient place to lie. The file command, the standard tool for this question on Linux and macOS, does not trust it. Its manual says there are three sets of tests, performed in this order: filesystem tests, magic tests, and language tests, and the first test that succeeds causes the file type to be printed. A magic test looks for an invariant identifier at a small fixed offset into the file.
Web browsers have the same problem with servers that label a file wrongly. The MIME Sniffing Standard exists because many HTTP servers supply a Content-Type that does not match the contents. It also describes why sniffing is risky: if a server believes the client will treat a file as an image, but the browser believes it to be HTML, an attacker might be able to steal the user's credentials. A file that is one thing in the name and another in its bytes is a problem in both directions.
What a signature looks like
Most formats start with a fixed signature. The PNG specification says the first eight bytes of a PNG datastream always contain the hexadecimal values 89 50 4E 47 0D 0A 1A 0A. A JPEG starts with FF D8 FF, and a GIF starts with GIF87a or GIF89a. A PDF starts with %PDF-, a ZIP file with PK, and a gzip file with 1F 8B. The page knows more than a hundred of them and shows the bytes it matched.
A Windows program is a good example of a signature with a second step. It starts with MZ, the DOS header. The Microsoft documentation says that at location 0x3c the stub holds the file offset to the PE signature, which is the four bytes PE followed by two zero bytes. A file that starts with MZ but has no PE signature is an old DOS program or a damaged file. The same documentation says a flag in the PE header marks the image as a dynamic-link library, which cannot be run directly, and the page uses it to tell a program from a library.
Containers: when the signature is not enough
Many formats are the same container with different contents. A Word document, an Excel workbook, a PowerPoint deck, an Android app, a Java archive, an EPUB book, and a plain ZIP archive all start with PK. The page reads the list of names that the ZIP file keeps at its end. According to the ZIP specification, a ZIP file must contain an end of central directory record, and a program finds the central directory by reading backward from the end of the file. A word/ folder means Word, an xl/ folder means Excel, a ppt/ folder means PowerPoint, and AndroidManifest.xml with a classes.dex means an Android app. A macro project inside makes it a file with macros.
OpenDocument files declare their type in a file named mimetype. The OpenDocument specification says that file shall be the first in the ZIP file, shall not be compressed, and holds the ASCII media type, so the page can read it without unpacking anything. Old Microsoft Office files are another container, and the page reads the names of the streams inside to tell a .doc from a .xls or a .ppt. MP4, MOV, HEIC, and AVIF files carry a brand in the first box, and Matroska and WebM files declare their DocType.
When the name and the contents disagree
The page gives each file one of four answers. It matches when the extension is one the type normally has. It is a note when the name is a related format, such as a .zip file that is really a Word document, or a name with no extension. It is a mismatch when the contents are a different type, such as a PNG named .jpg, and it says what to rename the file to. Renaming does not change the contents, and a PNG with a .jpg name usually still opens.
It is marked as unsafe when the contents are a program and the name says something else, such as invoice.pdf that is a Windows program. The page also reports tricks in the name itself: a double extension such as invoice.pdf.exe, in which only the last part counts, a direction-changing character that makes the end of a name display backwards, and long runs of spaces that push the extension out of view. None of these proves that a file is harmful. They mean the name is not telling you the truth, so the file deserves care.
Data attached to the end of a file
A well-formed file stops where its format says it stops. The PNG specification says the IEND chunk marks the end of the datastream and shall be last. A JPEG ends with a marker, a GIF with a trailer byte, a PDF with %%EOF, and a ZIP file with its end of central directory record. Programs ignore whatever comes after, so a program, an archive, or a script can be attached there and the file still opens as a picture, a document, or an archive.
The page finds where the real data ends and looks at what follows. A hidden program or archive is reported as something to check. Some extra data is normal: phones add a second picture, a depth map, or the video of a Motion Photo after a JPEG, and a few padding bytes are harmless. A file that is valid as two formats, such as a program followed by a ZIP archive, is also reported. That is how a self-extracting archive is made, and it is also how a picture is made to carry an archive.
Text files have no signature
A plain text file has nothing at its start that says what it is. The page reads the text instead. It checks the byte order mark and the encoding, including UTF-16 text without a mark. It recognizes an XML declaration, an HTML page, an SVG image, a script header such as #!/bin/sh, JSON that parses as a whole, JSON Lines, a PEM certificate, a vCard, an iCalendar, and tables with commas, tabs, semicolons, or pipes. A text that looks like JSON but does not parse is reported as text, with the error. Anything else is plain text, and any text extension is accepted for it.
How the checker was tested
The checker was run on 3,064 real files from one computer, from system folders, program folders, fonts, documents, and caches, across 513 extensions, and its answer was compared with libmagic 5.48, the library behind the file command. They gave the same answer for 93.4% of the files, after counting different spellings of one type, such as x-dosexec and vnd.microsoft.portable-executable, as the same. Most of the rest are types this page does not name, such as compiled terminfo entries and help files, and the many languages of source code that libmagic names and this page calls plain text. In some cases this page knew and libmagic 5.48 did not, for example WebAssembly modules and some fonts.
The checker was also run on 116 files made with other tools: Pillow, reportlab, python-docx, python-pptx, openpyxl, odfpy, ebooklib, and Python's tarfile, gzip, bz2, lzma, and zstandard, plus files built by hand to test headers, brands, and trailing data. It named the right type for every one and reported the findings that were planted: a ZIP attached to a PNG, a program attached to a JPEG, a PDF with text before its header, macros in a document, and a PNG, JPEG, ZIP, and PDF that were cut short. The file names in the real set were not looked at, only the answers. Of the 3,064 files, 10 had a name that did not match their contents. Three of them were .woff files that hold WOFF2 data.
Limits and accuracy
- The page identifies a file by its signature and structure. It does not scan for viruses and cannot tell a harmless program from a harmful one. A file marked as a program is only a program, not a threat.
- A format without a signature at its start, or one this page does not know, is reported as not recognized. More than a hundred types are known, which is fewer than libmagic's.
- A very large file is read at its first and last 4 MB. A ZIP file with a directory larger than that, or a trailing check inside the middle of a huge file, is not examined.
- An encrypted file shows as random data, or as a password-protected Office file, and its real type cannot be told without the password.
- Some answers are a likely guess, not a certainty, such as an MP3 without a tag, a table made of text, or a file without a clear signature. The page says how certain it is.
- A file that is cut off is reported only for the formats that have a clear end marker: PNG, JPEG, GIF, PDF, and ZIP.
- MIME types are the common names for each type. Servers and programs sometimes use others for the same type.
Frequently asked questions
How do I check the type of a file?
Add the file on this page. It reads the first bytes of the file, and looks inside ZIP, Office, and video containers when needed, then names the real type, shows its MIME type, and says whether the file name agrees. Nothing is uploaded.
What does it mean when the name does not match the contents?
The extension in the name is not one the real type uses, such as a PNG named .jpg. Often it was renamed or saved wrongly. If the contents are a program and the name is a document or a picture, the page marks it as unsafe, because that is a common way to disguise a program.
Can I trust a file that this page says is a PDF?
It means the file really is built as a PDF and has no obvious data attached. It does not mean the document is safe, because a PDF can contain links, scripts, or attachments. The page tells you only what the file is, not whether its contents are harmless.
Why does a Word file show as a ZIP file?
A .docx file is a ZIP archive that holds the parts of the document. The page looks at the names inside and tells a Word, Excel, or PowerPoint file from a plain ZIP. If a file named .docx shows as a plain ZIP, the parts of a Word document are missing and the file is damaged or not a document.
What is a magic number?
It is a fixed value at a fixed place in a file, usually at its start, that identifies the format. The PNG signature is 89 50 4E 47 0D 0A 1A 0A. Operating systems and the file command use these values to find the type of a file when the name cannot be trusted.
Why does a picture have extra data after its end?
Phones often store a second picture, a depth map, or a video after the main JPEG, and a few padding bytes are harmless. A program or an archive after the end of a picture is not normal, and the page marks it so you can check where the file came from.
Is my file uploaded?
No. Only the first and last parts of the file are read in your browser, and the code behind the page makes no network requests. The names and the results stay on your device and are gone when you close the page.
Research and references
This page was written and checked against the sources below.

