XML to Text
Extract the plain text content from an XML document.
XML to Text Converter Online
The XML to text converter strips the tags from an XML document and leaves you with the information inside it. Paste a feed, a configuration file, an export from a CMS, or any other XML, and the readable content appears immediately — ready to paste into an email, a ticket, a spreadsheet cell, or a search index.
How it works
The document is parsed with a strict XML parser, so the output follows the real element structure instead of a naive "remove everything between < and >" regex. That matters: entities like & become &, CDATA blocks keep their content, comments disappear, and attributes are never mistaken for text. The tree is then walked from the root element down, in document order.
Three output styles
- Text only – one line per element that contains text. This is the cleanest option for reading, word counting, or feeding content into a translation or spell-check tool.
- Path: value – every value is labelled with its full path, e.g.
catalog.book.title: Midnight Rain. Great for grepping, for comparing two exports line by line, or for documenting which fields a file contains. - Indented outline – a lightweight tree view as plain text, with two spaces per nesting level. Useful for pasting the structure of a document into notes or chat where no XML viewer is available.
Attributes
Switch Include attributes on to keep attribute values in the path and outline styles. Namespace declarations count as attributes too, so turn the option off if you only care about element content.
Typical use cases
Extracting article text from an RSS or Atom feed, pulling human-readable messages out of a SOAP fault, preparing XML content for full-text search, or quickly checking what data a partner actually sent you.
Related tools
To keep the structure instead of flattening it, convert with XML to JSON or XML to YAML. Tabular data works better with XML to CSV, and the XML Viewer shows the document as a collapsible tree.
Frequently Asked Questions
What is the difference between the three output styles?
“Text only” prints the text of every element on its own line and drops all markup. “Path: value” prefixes each value with its dotted element path, such as catalog.book.title. “Indented outline” lists every element name, indented by depth, followed by its text.
Are attribute values included?
In the path and outline styles, yes — as long as “Include attributes” is on. Paths show them as catalog.book@id: bk101 and the outline lists them in brackets after the element name. The text-only style never shows attributes, since it keeps just the readable content.
What happens to entities, CDATA, and comments?
Entities such as & and < are decoded into the characters they stand for, CDATA sections are output as plain text, and comments and processing instructions are skipped. Whitespace inside an element is collapsed to single spaces.
Why do I get an error instead of text?
The input has to be well-formed XML. If a tag is not closed or an attribute is missing its quotes, the tool reports the problem with its line and column. Fix it there, or check the document first with the XML Validator.