Parsing XML in Python with regex

You normally don't want to use re.match. Quoting from the docs:

If you want to locate a match anywhere in string, use search() instead (see also search() vs. match()).

Note:

>>> print re.match('>.*<', line)None>>> print re.search('>.*<', line)<_sre.SRE_Match object at 0x10f666238>>>> print re.search('>.*<', line).group(0)>PLAINSBORO, NJ 08536-1906<

Also, why parse XML with regex when you can use something like BeautifulSoup :).

>>> from bs4 import BeautifulSoup as BS>>> line='<City_State>PLAINSBORO, NJ 08536-1906</City_State>'>>> soup = BS(line)>>> print soup.find('city_state').textPLAINSBORO, NJ 08536-1906

python xml regex

Please, just use an XML parser like ElementTree

>>> from xml.etree import ElementTree as ET>>> line='<City_State>PLAINSBORO, NJ 08536-1906</City_State>'>>> ET.fromstring(line).text'PLAINSBORO, NJ 08536-1906'

python xml regex

re.match returns a match only if the pattern matches the entire string. To find substrings matching the pattern, use re.search.

And yes, this is a simple way to parse XML, but I would highly encourage you to use a library specifically designed for the task.

CodeHunter

Parsing XML in Python with regex

Recent Posts

How can I color dots in a xy scatterplot according to column value?

How to update a claim in ASP.NET Identity?

What does {0} mean when initializing an object?

Accessing members of items in a JSONArray with Java

How to log SQL statements in Spring Boot?

Powershell Get-WebSite name parameter is ignored

How to detect scroll to bottom of html element

Java synchronized method

How to test controllers with CodeIgniter?

Detect Visual Composer

Matplotlib: Specify format of floats for tick labels

Rails join a list of strings with commas and "and" before the last