Python查找文件中包含中文的行

最新推荐文章于 2024-07-28 20:22:30 发布

AlbertS

最新推荐文章于 2024-07-28 20:22:30 发布

阅读量6.9k

点赞数 1

分类专栏： Practical Python 文章标签： python utf-8 正则表达式 withopen 多语言版本

本文链接：https://blog.csdn.net/albertsh/article/details/78128042

版权

Practical 同时被 2 个专栏收录

60 篇文章 4 订阅

订阅专栏

Python

20 篇文章 3 订阅

订阅专栏

前言

近几天在做多语言版本的时候再次发现，区分各种语言真的是一件比较困难的事情，上一次做中文提取工具的就花了不少时间，这次决定用python试一试，结果写起来发现真是方便不少，自己整理了一下方便以后查找使用。

代码

#!/usr/bin/env python3
# -*- coding: utf-8 -*-
# find the line of containing chinese in files

__author__ = 'AlbertS'

import re

def start_find_chinese():
    find_count = 0;
    with open('ko_untranslated.txt', 'wb') as outfile:
        with open('source_ko.txt', 'rb') as infile:
            while True:
                content = infile.readline()
                if re.match(r'(.*[\u4E00-\u9FA5]+)|([\u4E00-\u9FA5]+.*)', content.decode('utf-8')):
                    outfile.write(content)
                    find_count += 1;

                if not content:
                    return find_count

# start to find
if __name__ == '__main__':
    count = start_find_chinese()
    print("find complete! count =", count)

原始文件

source_ko.txt文件内容

3   캐릭터 Lv.50 달성
8   캐릭터 Lv.80 달성
10  캐릭터 Lv.90 달성
...
...
2840    飞黄腾达
4841    同归于尽
8848    캐릭터 Lv.50 달

运行效果(ko_untranslated.txt文件)

2840    飞黄腾达
4841    同归于尽

总结

其实这段小小的代码中包含了两个常用的功能，那就是读写文件和正则表达式。
这也是两个重要的知识点，其中with操作可能防止资源泄漏，操作起来更加方便。
正则表达式可是一个文字处理的利器，代码中的正则可能还不太完善，后续我会继续补充更新。

AlbertS

关注

1
点赞
踩
7

收藏

觉得还不错? 一键收藏
打赏
0
评论
复制链接

分享到 QQ

分享到新浪微博

扫一扫

专栏目录