什么是XML

XML是可扩展标记语言（Extensible Markup Language）的缩写，其中标记是关键部分。用户可以创建内容，然后使用限定标记标记它，从而使每个单词、短语或块成为可识别、可分类的信息。
标记语言从早起的私有公司和政府制定形式逐渐演变成标准通用标记语言（Standard Generalized Markup Language，SGML）、超文本标记语言（Hypertext Markup Language，HTML），并且最终演变成XML。XML有以下几个特点：

XML的设计宗旨是传输数据，而非显示数据

XML的标签没有被预定义，用户需要自行定义标签

XML被设计为具有自我描述性

XML是W3C的推荐标准

Python对XML文件的解析

常见的XML编程接口有DOM和SAX，这两种接口处理XML文件的方式不同，使用场合也不同。DOM是由W3C官方提出的标准，它会把整个XML文件读入内存，并将该文件解析成树，我们可以通过访问树的节点的方式访问XML中的标签，但是这种方法占用内存大，解析慢，如果读入文件过大，尽量避免使用这种方法。SAX是事件驱动的，通过在解析XML的过程中触发一个个的事件并调用用户自定义的回调函数来处理XML文件，速度比较快，占用内存少，但是需要用户实现回调函数，因此Python标准库的官方文档中这样介绍SAX：SAX每次只允许你查看文档的一小部分，你无法通过当前获取的元素访问其他元素。Python中提供了很多包支持XML文件的解析，如xml.dom，xml.sax，xml.dom.minidom和xml.etree.ElementTree等，本文重点介绍xml.dom.minidom。

xml.dom.minidom包

xml.dom.minidom是DOM API的极简化实现，比完整版的DOM要简单的多，而且这个包也小得多，下面以movie.xml文件为例进行操作。

<collection shelf="New Arrivals">

<movie title="Enemy Behind">

   <type>War, Thriller</type>

   <format>DVD</format>

   <year>2003</year>

   <rating>PG</rating>

   <stars>10</stars>

   <description>Talk about a US-Japan war</description>

</movie>

<movie title="Transformers">

   <type>Anime, Science Fiction</type>

   <format>DVD</format>

   <year>1989</year>

   <rating>R</rating>

   <stars>8</stars>

   <description>A schientific fiction</description>

</movie>

   <movie title="Trigun">

   <type>Anime, Action</type>

   <format>DVD</format>

   <episodes>4</episodes>

   <rating>PG</rating>

   <stars>10</stars>

   <description>Vash the Stampede!</description>

</movie>

<movie title="Ishtar">

   <type>Comedy</type>

   <format>VHS</format>

   <rating>PG</rating>

   <stars>2</stars>

   <description>Viewable boredom</description>

</movie>

</collection>

然后我们调用xml.dom.minidom.parse方法读入xml文件并解析成DOM树

from xml.dom.minidom import parse

import xml.dom.minidom

# 使用minidom解析器打开 XML 文档

DOMTree = xml.dom.minidom.parse("F:/project/Breast/codes/AllXML/aa.xml")

collection = DOMTree.documentElement

if collection.hasAttribute("shelf"):

    print("Root element : %s" % collection.getAttribute("shelf"))

# 在集合中获取所有电影

movies = collection.getElementsByTagName("movie")

# 打印每部电影的详细信息

for movie in movies:

    print("*****Movie*****")

    if movie.hasAttribute("title"):

        print("Title: %s" % movie.getAttribute("title"))

    type = movie.getElementsByTagName('type')[0]

    print("Type: %s" % type.childNodes[0].data)

    format = movie.getElementsByTagName('format')[0]

    print("Format: %s" % format.childNodes[0].data)

    rating = movie.getElementsByTagName('rating')[0]

    print("Rating: %s" % rating.childNodes[0].data)

    description = movie.getElementsByTagName('description')[0]

    print("Description: %s" % description.childNodes[0].data)

以上程序执行结果如下：

Root element : New Arrivals

*****Movie*****

Title: Enemy Behind

Type: War, Thriller

Format: DVD

Rating: PG

Description: Talk about a US-Japan war

*****Movie*****

Title: Transformers

Type: Anime, Science Fiction

Format: DVD

Rating: R

Description: A schientific fiction

*****Movie*****

Title: Trigun

Type: Anime, Action

Format: DVD

Rating: PG

Description: Vash the Stampede!

*****Movie*****

Title: Ishtar

Type: Comedy

Format: VHS

Rating: PG

Description: Viewable boredom

实战—批量修改XML文件

最近在用caffe-ssd训练比赛的数据集，但是官方给的数据集用来标记的XML文件并不是标准格式，一些标签的命名不对，导致无法正确生成lmdb文件，因此需要修改这些标签，下面用Python实现了一个批量修改XML文件的脚本。

# -*- coding:utf-8 -*-

import os

import xml.dom.minidom

xml_file_path = "/home/lyz/data/VOCdevkit/MyDataSet/Annotations/"

lst_label = ["height", "width", "depth"]

lst_dir = os.listdir(xml_file_path)

for file_name in lst_dir:

    file_path = xml_file_path + file_name

    tree = xml.dom.minidom.parse(file_path)

    root = tree.documentElement        #获取根结点

    size_node = root.getElementsByTagName("size")[0]

    for size_label in lst_label:    #替换size标签下的子节点

        child_tag = "img_" + size_label

        child_node = size_node.getElementsByTagName(child_tag)[0]

        new_node = tree.createElement(size_label)

        text = tree.createTextNode(child_node.firstChild.data)

        new_node.appendChild(text)

        size_node.replaceChild(new_node, child_node)

    #替换object下的boundingbox节点

    lst_obj = root.getElementsByTagName("object")

    data = {}

    for obj_node in lst_obj:

        box_node = obj_node.getElementsByTagName("bounding_box")[0]

        new_box_node = tree.createElement("bndbox")

        for child_node in box_node.childNodes:

            tmp_node = child_node.cloneNode("deep")

            new_box_node.appendChild(tmp_node)

        x_node = new_box_node.getElementsByTagName("x_left_top")[0]

        xmin = x_node.firstChild.data

        data["xmin"] = (xmin, x_node)

        y_node = new_box_node.getElementsByTagName("y_left_top")[0]

        ymin = y_node.firstChild.data

        data["ymin"] = (ymin, y_node)

        w_node = new_box_node.getElementsByTagName("width")[0]

        xmax = str(int(xmin) + int(w_node.firstChild.data))

        data["xmax"] = (xmax, w_node)

        h_node = new_box_node.getElementsByTagName("height")[0]

        ymax = str(int(ymin) + int(h_node.firstChild.data))

        data["ymax"] = (ymax, h_node)

        for k, v in data.items():

            new_node = tree.createElement(k)

            text = tree.createTextNode(v[0])

            new_node.appendChild(text)

            new_box_node.replaceChild(new_node, v[1])

        obj_node.replaceChild(new_box_node, box_node)

    with open(file_path, 'w') as f:

        tree.writexml(f, indent="\n", addindent="\t", encoding='utf-8')

    #去掉XML文件头（一些情况下文件头的存在可能导致错误）

    lines = []

    with open(file_path, 'rb') as f:

        lines = f.readlines()[1:]

    with open(file_path, 'wb') as f:

        f.writelines(lines)

print("-----------------done--------------------")

关于writexml方法：indent参数表示在当前节点之前插入的字符，addindent表示在该结点的子节点前插入的字符

读取只包含标签的xml的更多相关文章

读取配置文件包含properties和xml文件
读取properties配置文件 /** * 读取配置文件 * @author ll-t150 */ public class Utils { private static Properties pr ...
死磕Spring之IoC篇 - 解析自定义标签（XML 文件）
该系列文章是本人在学习 Spring 的过程中总结下来的,里面涉及到相关源码,可能对读者不太友好,请结合我的源码注释 Spring 源码分析 GitHub 地址进行阅读 Spring 版本:5.1. ...
19 标签：xml或者html
1 标签:xml或者html 1.1 使用XmlSlurper解析xml groovy处理xml非常容易.XmlSlurper 类用来处理xml.在处理xml方面,还有其他的处理方式,但 ...
51Nod 1010 只包含因子2 3 5的数 Label：None
K的因子中只包含2 3 5.满足条件的前10个数是:2,3,4,5,6,8,9,10,12,15. 所有这样的K组成了一个序列S,现在给出一个数n,求S中 >= 给定数的最小的数. 例如:n = ...
51Nod--1010 只包含235的数
51Nod: http://www.51nod.com/onlineJudge/questionCode.html#!problemId=1010 1010 只包含因子2 3 5的数基准时间限制:1 ...
做参数可以读取参数保存参数用xml文件的方式
做参数可以读取参数保存参数用xml文件的方式好处:供不同用户保存适合自己使用的参数
sql语句读取所有父子标签
select A.HOSPITAL_ID from T_HOSPITAL A connect by prior A.HOSPITAL_ID=A.PARENT_ID start with A.HOSPI ...
只包含schema的dll生成和引用方法
工作中,所有的tools里有一个project是只包含若干个schema的工程,研究了一下,发现创建这种只包含schema的dll其实非常简单. 首先,在visual studio-new proje ...
1007 正整数分组 1010 只包含因子2 3 5的数 1014 X^2 Mod P 1024 矩阵中不重复的元素 1031 骨牌覆盖
1007 正整数分组将一堆正整数分为2组,要求2组的和相差最小. 例如:1 2 3 4 5,将1 2 4分为1组,3 5分为1组,两组和相差1,是所有方案中相差最少的. Input 第1行:一个 ...

随机推荐

sed命令总结
目录 1.概述 2.查 1.打印整行(一或多) 2.正则打印包含关键字的行 2.增 3.删 4.改 5.后向引用 6.结合 7.练习我叫张贺,贪财好色.一名合格的LINUX运维工程师,专注于LINU ...
【未完成】【oracle】单引号使用问题
‘-’不可以用原因:
构建LVS负载均衡集群——NAT模式（最简单方式）
一.装备一台lvs调度器主机要求两个网卡一个为内部局域网ip,一个为公网ip #IP地址设置过程不再重复 [root@localhost ~]# ip a | grep eth0 #内网ip : et ...
Codeforces Round #594 (Div. 1) D2. The World Is Just a Programming Task (Hard Version) 括号序列思维
D2. The World Is Just a Programming Task (Hard Version) This is a harder version of the problem. In ...
celery生产者-消费者
Celery是一个简单,灵活,可靠的分布式系统,用于处理大量消息,同时为操作提供维护此类系统所需的工具. 它是一个任务队列,专注于实时处理,同时还支持任务调度. celery解决了什么问题: 示例一: ...
ReactNative: 使用View组件创建九宫格
一.简言初学RN,一切皆新.View组件跟我们iOS中UIView类似,作为一个容器视图使用,它主要负责承载其他的子组件.View组件采用的是FlexBox伸缩盒子布局,通过对它的布局可以影响子组件 ...
dva+umi+antd项目从搭建到使用（没有剖验证，不知道在说i什么）
先创建一个新项目,具体步骤请参考https://www.cnblogs.com/darkbluelove/p/11338309.html 一.添加document.ejs文件(参考文档:https:/ ...
抓包工具之fiddler实战3-接口测试
Fiddler实现接口测试 Fiddler提供了进行接口测试的功能,找到composer界面,选择接口方法,填写接口URL地址,发送请求. 例子:全国天气预报的接口 http://v.juhe.cn/ ...
Kubernetes V1.15 二进制部署集群
1. 架构篇 1.1 kubernetes 架构说明 1.2 Flannel网络架构图 1.3 Kubernetes工作流程 2. 组件介绍 2.1 ...
[04]ASP.NET Core Web 项目文件
ASP.NET Core Web 项目文件本文作者:梁桐铭- 微软最有价值专家(Microsoft MVP) 文章会随着版本进行更新,关注我获取最新版本本文出自<从零开始学 ASP.NET ...