国产探花免费观看_亚洲丰满少妇自慰呻吟_97日韩有码在线_资源在线日韩欧美_一区二区精品毛片,辰东完美世界有声小说,欢乐颂第一季,yy玄幻小说排行榜完本

首頁 > 編程 > Python > 正文

利用Python腳本生成sitemap.xml的實現(xiàn)方法

2019-11-25 16:22:46
字體:
來源:轉載
供稿:網(wǎng)友

安裝lxml

首先需要pip install lxml安裝lxml庫。

如果你在ubuntu上遇到了以下錯誤:

#include "libxml/xmlversion.h"compilation terminated.error: command 'x86_64-linux-gnu-gcc' failed with exit status 1----------------------------------------Cleaning up... Removing temporary dir /tmp/pip_build_root...Command /usr/bin/python -c "import setuptools, tokenize;__file__='/tmp/pip_build_root/lxml/setup.py';exec(compile(getattr(tokenize, 'open', open)(__file__).read().replace('/r/n', '/n'), __file__, 'exec'))" install --record /tmp/pip-O4cIn6-record/install-record.txt --single-version-externally-managed --compile failed with error code 1 in /tmp/pip_build_root/lxmlException information:Traceback (most recent call last): File "/usr/lib/python2.7/dist-packages/pip/basecommand.py", line 122, in main  status = self.run(options, args) File "/usr/lib/python2.7/dist-packages/pip/commands/install.py", line 283, in run  requirement_set.install(install_options, global_options, root=options.root_path) File "/usr/lib/python2.7/dist-packages/pip/req.py", line 1435, in install  requirement.install(install_options, global_options, *args, **kwargs) File "/usr/lib/python2.7/dist-packages/pip/req.py", line 706, in install  cwd=self.source_dir, filter_stdout=self._filter_install, show_stdout=False) File "/usr/lib/python2.7/dist-packages/pip/util.py", line 697, in call_subprocess  % (command_desc, proc.returncode, cwd))InstallationError: Command /usr/bin/python -c "import setuptools, tokenize;__file__='/tmp/pip_build_root/lxml/setup.py';exec(compile(getattr(tokenize, 'open', open)(__file__).read().replace('/r/n', '/n'), __file__, 'exec'))" install --record /tmp/pip-O4cIn6-record/install-record.txt --single-version-externally-managed --compile failed with error code 1 in /tmp/pip_build_root/lxml

請安裝以下依賴:

sudo apt-get install libxml2-dev libxslt1-dev

Python代碼

下面是生成sitemap和sitemapindex索引的代碼,可以按照需求傳入需要的參數(shù),或者增加字段:

#!/usr/bin/env python# -*- coding:utf-8 -*-import ioimport refrom lxml import etreedef generate_xml(filename, url_list):  """Generate a new xml file use url_list"""  root = etree.Element('urlset',             xmlns="http://www.sitemaps.org/schemas/sitemap/0.9")  for each in url_list:    url = etree.Element('url')    loc = etree.Element('loc')    loc.text = each    url.append(loc)    root.append(url)  header = u'<?xml version="1.0" encoding="UTF-8"?>/n'  s = etree.tostring(root, encoding='utf-8', pretty_print=True)  with io.open(filename, 'w', encoding='utf-8') as f:    f.write(unicode(header+s))def update_xml(filename, url_list):  """Add new url_list to origin xml file."""  f = open(filename, 'r')  lines = [i.strip() for i in f.readlines()]  f.close()  old_url_list = []  for each_line in lines:    d = re.findall('<loc>(http:////.+)<//loc>', each_line)    old_url_list += d  url_list += old_url_list  generate_xml(filename, url_list)def generatr_xml_index(filename, sitemap_list, lastmod_list):  """Generate sitemap index xml file."""  root = etree.Element('sitemapindex',             xmlns="http://www.sitemaps.org/schemas/sitemap/0.9")  for each_sitemap, each_lastmod in zip(sitemap_list, lastmod_list):    sitemap = etree.Element('sitemap')    loc = etree.Element('loc')    loc.text = each_sitemap    lastmod = etree.Element('lastmod')    lastmod.text = each_lastmod    sitemap.append(loc)    sitemap.append(lastmod)    root.append(sitemap)  header = u'<?xml version="1.0" encoding="UTF-8"?>/n'  s = etree.tostring(root, encoding='utf-8', pretty_print=True)  with io.open(filename, 'w', encoding='utf-8') as f:    f.write(unicode(header+s))if __name__ == '__main__':  urls = ['http://www.baidu.com'] * 10  mods = ['2004-10-01T18:23:17+00:00'] * 10  generatr_xml_index('index.xml', urls, mods)

效果

生成的效果應該是這種格式:

sitemap格式:

<?xml version="1.0" encoding="UTF-8"?><urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9"> <url>  <loc>http://www.example.com/foo.html</loc> </url></urlset>

sitemapindex格式:

<?xml version="1.0" encoding="UTF-8"?>  <sitemapindex xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">  <sitemap>   <loc>http://www.example.com/sitemap1.xml.gz</loc>   <lastmod>2004-10-01T18:23:17+00:00</lastmod>  </sitemap>  <sitemap>   <loc>http://www.example.com/sitemap2.xml.gz</loc>   <lastmod>2005-01-01</lastmod>  </sitemap>  </sitemapindex>

lastmod時間格式的問題

格式是用ISO 8601的標準,如果是linux/unix系統(tǒng),可以使用以下函數(shù)獲取

def get_lastmod_time(filename):  time_stamp = os.path.getmtime(filename)  t = time.localtime(time_stamp)  # return time.strftime('%Y-%m-%dT%H:%M:%S+08:00', t)  return time.strftime('%Y-%m-%dT%H:%M:%SZ', t)

優(yōu)化

一般來說,用lxml效率低并且內(nèi)存占用比較大,可以直接用文件的write方法創(chuàng)建。

def generate_xml(filename, url_list):  with gzip.open(filename,"w") as f:    f.write("""<?xml version="1.0" encoding="utf-8"?><urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">/n""")    for i in url_list:      f.write("""<url><loc>%s</loc></url>/n"""%i)    f.write("""</urlset>""")def append_xml(filename, url_list):  with gzip.open(filename, 'r') as f:    for each_line in f:      d = re.findall('<loc>(http:////.+)<//loc>', each_line)      url_list.extend(d)    generate_xml(filename, set(url_list))def modify_time(filename):  time_stamp = os.path.getmtime(filename)  t = time.localtime(time_stamp)  return time.strftime('%Y-%m-%dT%H:%M:%S:%SZ', t)def new_xml(filename, url_list):  generate_xml(filename, url_list)  root = dirname(filename)  with open(join(dirname(root), "sitemap.xml"),"w") as f:    f.write('<?xml version="1.0" encoding="utf-8"?>/n<sitemapindex xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">/n')    for i in glob.glob(join(root,"*.xml.gz")):      lastmod = modify_time(i)      i = i[len(CONFIG.SITEMAP_PATH):]      f.write("<sitemap>/n<loc>http:/%s</loc>/n"%i)      f.write("<lastmod>%s</lastmod>/n</sitemap>/n"%lastmod)    f.write('</sitemapindex>')

總結

以上就是這篇文章的全部內(nèi)容了,希望本文的內(nèi)容對大家學習或者使用python能帶來一定的幫助,如果有疑問大家可以留言交流。謝謝大家對武林網(wǎng)的支持。

發(fā)表評論 共有條評論
用戶名: 密碼:
驗證碼: 匿名發(fā)表
主站蜘蛛池模板: 蒙山县| 句容市| 武平县| 卢湾区| 长汀县| 康平县| 涪陵区| 天祝| 夏津县| 资兴市| 清涧县| 家居| 吉林市| 高唐县| 神池县| 明星| 赫章县| 临夏县| 咸宁市| 汶上县| 普兰店市| 龙门县| 高邮市| 金昌市| 大冶市| 威远县| 房山区| 尚义县| 宣化县| 三河市| 山阳县| 封开县| 通山县| 抚松县| 汾阳市| 仁化县| 浮山县| 长治市| 贡山| 天长市| 新兴县|