python实现获取文件列表中每一个文件keyword

功能描写叙述：

获取某个路径下的全部文件，提取出每一个文件里出现频率最高的前300个字。保存在数据库其中。

前提。你须要配置好nltk

#!/usr/bin/python

#coding=utf-8

'''

function : This script will create a database named mydb then

           abstract keywords of files of privacy police.

author    : Chicho

date      : 2014/7/28

running   : python key_extract.py -d path_of_file

'''

import sys,getopt

import nltk

import MySQLdb

from nltk.corpus import PlaintextCorpusReader

corpus_root = ""

if __name__ == '__main__':

    opts,args = getopt.getopt(sys.argv[1:], "d:h","directory=help")

    #get the directory

    for op,value in opts:

        if op in ("-d", "--directory"):

            corpus_root = value

	#actually。 the above method to get  a directory is a little complicated,you can

	#do like this

	'''

	the input include you path and use sys.argv to get the path

	'''

	'''

	running : python key_extract.py you path_of_file

	corpus_root = sys.argv[1]

	'''

    # corpus_root is the directory of files of privacy policy, all of the are html files

    filelists = PlaintextCorpusReader(corpus_root, '.*')

    #get the files' list

    files = filelists.fileids()

    #connect the database

    conn = MySQLdb.connect(host = 'your_personal_host_ip_address', user = 'rusername', port =your_port, passwd = 'U_password')

    #get the cursor

    curs = conn.cursor()

    conn.set_character_set('utf8')

    curs.execute('set names utf8')

    curs.execute('SET CHARACTER SET utf8;')

    curs.execute('SET character_set_connection=utf8;')

    '''

    conn.text_factory=lambda x: unicode(x, 'utf8', "ignore")

    #conn.text_factory=str

    '''	 

    # create a database named mydb

    '''

    try:

        curs.execute("create database mydb")

    except Exception,e:

        print e

    '''

    conn.select_db('mydb')

    try:

        for i in range(300):

            sql = "alter table filekeywords add " + "key" + str(i) + " varchar(45)"

            curs.execute(sql)

    except Exception,e:

        print e

    i = 0

    for privacyfile in files:

        #f = open(privacyfile,'r', encoding= 'utf-8')

        sql = "insert into filekeywords set id =" + str(i)

        curs.execute(sql)

        sql = "update filekeywords set name =" + "'" + privacyfile + "' where id= " + str(i)

        curs.execute(sql)

        # get the words in privacy policy

        wordlist = [w for w in filelists.words(privacyfile) if w.isalpha() and len(w)>2]

        # get the keywords

        fdist = nltk.FreqDist(wordlist)

        vol = fdist.keys()

        key_num = len(vol)

        if key_num > 300:

            key_num = 300

        for j in range(key_num):

            sql = "update filekeywords set " + "key" + str(j) + "=" + "'" + vol[j] + "' where id=" + str(i)

            curs.execute(sql)

        i = i + 1

    conn.commit()

    curs.close()

    conn.close()

转载注明出处：http://blog.csdn.net/chichoxian/article/details/42003603

python实现获取文件列表中每一个文件keyword的更多相关文章

java中的文件读取和文件写出：如何从一个文件中获取内容以及如何向一个文件中写入内容
import java.io.BufferedReader; import java.io.BufferedWriter; import java.io.File; import java.io.Fi ...
python实现获取文件夹中的最新文件
实现代码如下: #查找某目录中的最新文件import osclass FindNewFile: def find_NewFile(self,path): #获取文件夹中的所有文件 lists = os ...
基于Python——实现解压文件夹中的.zip文件
[背景]当一个文件夹里存好好多.zip文件需要解压时,手动一个个解压再给文件重命名是一件很麻烦的事情,基于此,今天介绍一种使用python实现批量解压文件夹中的压缩文件并给文件重命名的方法—— [代码 ...
每日学习心得：SharePoint 为列表中的文件夹添加子项（文件夹）、新增指定内容类型的子项、查询列表中指定的文件夹下的内容
前言: 这里主要是针对列表中的文件下新增子项的操作,同时在新建子项时,可以为子项指定特定的内容类型,在某些时候需要查询指定的文件夹下的内容,针对这些场景都一一给力示例和说明,都是一些很小的知识点,希望 ...
python操作txt文件中数据教程[3]-python读取文件夹中所有txt文件并将数据转为csv文件
python操作txt文件中数据教程[3]-python读取文件夹中所有txt文件并将数据转为csv文件觉得有用的话,欢迎一起讨论相互学习~Follow Me 参考文献 python操作txt文件中 ...
python 将指定文件夹中的指定文件放入指定文件夹中
import os import shutil import re #获取指定文件中文件名 def get_filename(filetype): name =[] final_name_list = ...
在/proc文件系统中增加一个目录hello，并在这个目录中增加一个文件world，文件的内容为hello world
一.题目编写一个内核模块,在/proc文件系统中增加一个目录hello,并在这个目录中增加一个文件world,文件的内容为hello world.内核版本要求2.6.18 二.实验环境物理主机:w ...
获取SD卡中的音乐文件
小编近期在搞一个音乐播放器App.练练手: 首先遇到一个问题.怎么获取本地的音乐文件? /** * 获取SD卡中的音乐文件 * * @param context * @return */ public ...
创建一个目录info,并在目录中创建一个文件test.txt,把该文件的信息读取出来，并显示出来
/*4.创建一个目录info,并在目录中创建一个文件test.txt,把该文件的信息读取出来,并显示出来*/ #import <Foundation/Foundation.h>#defin ...

随机推荐

【Henu ACM Round#19 C】 Developing Skills
[链接] 我是链接,点我呀:) [题意] 在这里输入题意 [题解] 优先把不是10的倍数的变成10的倍数. (优先%10比较大的数字增加如果k还有剩余. 剩下的数字都是10的倍数了. 那么先加哪一个 ...
JAVA利用反射映射JSON对象为JavaBean
关于将JSONObject转换为JavaBean,其实在JSONObject中有对于的toBean()方法来处理,还可以根据给定的JsonConfig来处理一些相应的要求,比如过滤指定的属性 //返回 ...
3D打印技术之切片引擎（6）
[此系列文章基于熔融沉积( fused depostion modeling, FDM )成形工艺] 这一篇文章说一下填充算法中的网格填充.网格填充在现有的较为成熟的引擎中是非常普遍的:skeinfo ...
hdu 1102 Constructing Roads(kruskal || prim)
求最小生成树.有一点点的变化,就是有的边已经给出来了.所以,最小生成树里面必须有这些边,kruskal和prim算法都能够,prim更简单一些.有一点须要注意,用克鲁斯卡尔算法的时候须要将已经存在的边 ...
HotSpotVM的Java堆实现浅析#1：初始化
今天来看看HotSpotVM的Java堆初始化. Universe Java堆的初始化主要由Universe模块来完毕,来看下Universe模块初始化的代码,universe_init. jint ...
iBatis框架使用 4步曲
iBatis是一款使用方便的数据訪问工具,也可作为数据持久层的框架.和ORM框架(如Hibernate)将数据库表直接映射为Java对象相比.iBatis是将SQL语句映射为Java对象. 相对于全自 ...
swust oj 2516 教练我想学算术 dp+组合计数
#include<stdio.h> #include<string.h> #include<iostream> #include<string> #in ...
poj--2007--Scrambled Polygon(数学几何基础)
Scrambled Polygon Time Limit: 1000MS Memory Limit: 30000KB 64bit IO Format: %I64d & %I64u Su ...
Linq中where查询
Linq的Where操作包括3种形式:简单形式.关系条件形式.First()形式. 1.简单形式: 例:使用where查询在北京的客户 var q = from c in db.Customers ...
ThinkPad X260 UEFI安装 win7 64位方法
ThinkPad X260 UEFI安装 win7 64位方法 1.使用DG重新格式化硬盘,格式为GPT 2.使用CGI 安装 WIM文件 (image不知是否可以,下次测试) 3.改BIOS ...

python实现获取文件列表中每一个文件keyword

python实现获取文件列表中每一个文件keyword的更多相关文章

随机推荐

热门专题