python讀取一個大於10G的txt文件的方法

前言

用python 讀取一個大於10G 的文件,自己電腦隻有8G內存,一運行就報內存溢出:MemoryError
python 如何用open函數讀取大文件呢?

讀取大文件

首先可以自己先制作一個大於10G的txt文件

a = '''
2021-02-02 21:33:31,678 [django.request:93] [base:get_response] [WARNING]- Not Found: /http:/123.125.114.144/
2021-02-02 21:33:31,679 [django.server:124] [basehttp:log_message] [WARNING]- "HEAD http://123.125.114.144/ HTTP/1.1" 404 1678
2021-02-02 22:14:04,121 [django.server:124] [basehttp:log_message] [INFO]- code 400, message Bad request version ('HTTP')
2021-02-02 22:14:04,122 [django.server:124] [basehttp:log_message] [WARNING]- "GET ../../mnt/custom/ProductDefinition HTTP" 400 -
2021-02-02 22:16:21,052 [django.server:124] [basehttp:log_message] [INFO]- "GET /api/login HTTP/1.1" 301 0
2021-02-02 22:16:21,123 [django.server:124] [basehttp:log_message] [INFO]- "GET /api/login/ HTTP/1.1" 200 3876
2021-02-02 22:16:21,192 [django.server:124] [basehttp:log_message] [INFO]- "GET /static/assets/img/main_bg.png HTTP/1.1" 200 2801
2021-02-02 22:16:21,196 [django.server:124] [basehttp:log_message] [INFO]- "GET /static/assets/iconfont/style.css HTTP/1.1" 200 1638
2021-02-02 22:16:21,229 [django.server:124] [basehttp:log_message] [INFO]- "GET /static/assets/img/bg.jpg HTTP/1.1" 200 135990
2021-02-02 22:16:21,307 [django.server:124] [basehttp:log_message] [INFO]- "GET /static/assets/iconfont/fonts/icomoon.ttf?u4m6fy HTTP/1.1" 200 6900
2021-02-02 22:16:23,525 [django.server:124] [basehttp:log_message] [INFO]- "POST /api/login/ HTTP/1.1" 302 0
2021-02-02 22:16:23,618 [django.server:124] [basehttp:log_message] [INFO]- "GET /api/index/ HTTP/1.1" 200 18447
2021-02-02 22:16:23,709 [django.server:124] [basehttp:log_message] [INFO]- "GET /static/assets/js/commons.js HTTP/1.1" 200 13209
2021-02-02 22:16:23,712 [django.server:124] [basehttp:log_message] [INFO]- "GET /static/assets/css/admin.css HTTP/1.1" 200 19660
2021-02-02 22:16:23,712 [django.server:124] [basehttp:log_message] [INFO]- "GET /static/assets/css/common.css HTTP/1.1" 200 1004
2021-02-02 22:16:23,714 [django.server:124] [basehttp:log_message] [INFO]- "GET /static/assets/js/app.js HTTP/1.1" 200 20844
2021-02-02 22:16:26,509 [django.server:124] [basehttp:log_message] [INFO]- "GET /api/report_list/1/ HTTP/1.1" 200 14649
2021-02-02 22:16:51,496 [django.server:124] [basehttp:log_message] [INFO]- "GET /api/test_list/1/ HTTP/1.1" 200 24874
2021-02-02 22:16:51,721 [django.server:124] [basehttp:log_message] [INFO]- "POST /api/add_case/ HTTP/1.1" 200 0
2021-02-02 22:16:59,707 [django.server:124] [basehttp:log_message] [INFO]- "GET /api/test_list/1/ HTTP/1.1" 200 24874
2021-02-03 22:16:59,909 [django.server:124] [basehttp:log_message] [INFO]- "POST /api/add_case/ HTTP/1.1" 200 0
2021-02-03 22:17:01,306 [django.server:124] [basehttp:log_message] [INFO]- "GET /api/edit_case/1/ HTTP/1.1" 200 36504
2021-02-03 22:17:06,265 [django.server:124] [basehttp:log_message] [INFO]- "GET /api/add_project/ HTTP/1.1" 200 17737
2021-02-03 22:17:07,825 [django.server:124] [basehttp:log_message] [INFO]- "GET /api/project_list/1/ HTTP/1.1" 200 29789
2021-02-03 22:17:13,116 [django.server:124] [basehttp:log_message] [INFO]- "GET /api/add_config/ HTTP/1.1" 200 24816
2021-02-03 22:17:19,671 [django.server:124] [basehttp:log_message] [INFO]- "GET /api/config_list/1/ HTTP/1.1" 200 19532
'''
while True:
    with open("xxx.log", "a", encoding="utf-8") as fp:
         fp.write(a)

循環寫入到 xxx.log 文件,運行 3-5 分鐘,pycharm 打開查看文件大小大於 10G

於是我用open函數 直接讀取

f = open("xxx.log", 'r')
print(f.read())
f.close()

拋出內存溢出異常:MemoryError

Traceback (most recent call last):
File “D:/2021kecheng06/demo/txt.py”, line 35, in <module>
print(f.read())
MemoryError

運行的時候可以看下自己電腦的內存已經占瞭100%, cpu高達91% ,不掛掉才怪瞭!

這種錯誤的原因在於,read()方法執行操作是一次性的都讀入內存中,顯然文件大於內存就會報錯。

read() 的幾種方法

1.read() 方法可以帶參數 n, n 是每次讀取的大小長度,也就是可以每次讀一部分,這樣就不會導致內存溢出

f = open("xxx.log", 'r')
print(f.read(2048))
f.close()

運行結果

2019-10-24 21:33:31,678 [django.request:93] [base:get_response] [WARNING]- Not Found: /http:/123.125.114.144/
2019-10-24 21:33:31,679 [django.server:124] [basehttp:log_message] [WARNING]- “HEAD http://123.125.114.144/ HTTP/1.1” 404 1678
2019-10-24 22:14:04,121 [django.server:124] [basehttp:log_message] [INFO]- code 400, message Bad request version (‘HTTP’)
2019-10-24 22:14:04,122 [django.server:124] [basehttp:log_message] [WARNING]- “GET ../../mnt/custom/ProductDefinition HTTP” 400 –
2019-10-24 22:16:21,052 [django.server:124] [basehttp:log_message] [INFO]- “GET /api/login HTTP/1.1” 301 0
2019-10-24 22:16:21,123 [django.server:124] [basehttp:log_message] [INFO]- “GET /api/login/ HTTP/1.1” 200 3876
2019-10-24 22:16:21,192 [django.server:124] [basehttp:log_message] [INFO]- “GET /static/assets/img/main_bg.png HTTP/1.1” 200 2801
2019-10-24 22:16:21,196 [django.server:124] [basehttp:log_message] [INFO]- “GET /static/assets/iconfont/style.css HTTP/1.1” 200 1638
2019-10-24 22:16:21,229 [django.server:124] [basehttp:log_message] [INFO]- “GET /static/assets/img/bg.jpg HTTP/1.1” 200 135990
2019-10-24 22:16:21,307 [django.server:124] [basehttp:log_message] [INFO]- “GET /static/assets/iconfont/fonts/icomoon.ttf?u4m6fy HTTP/1.1” 200 6900
2019-10-24 22:16:23,525 [django.server:124] [basehttp:log_message] [INFO]- “POST /api/login/ HTTP/1.1” 302 0
2019-10-24 22:16:23,618 [django.server:124] [basehttp:log_message] [INFO]- “GET /api/index/ HTTP/1.1” 200 18447
2019-10-24 22:16:23,709 [django.server:124] [basehttp:log_message] [INFO]- “GET /static/assets/js/commons.js HTTP/1.1” 200 13209
2019-10-24 22:16:23,712 [django.server:124] [basehttp:log_message] [INFO]- “GET /static/assets/css/admin.css HTTP/1.1” 200 19660
2019-10-24 22:16:23,712 [django.server:124] [basehttp:log_message] [INFO]- “GET /static/assets/css/common.css HTTP/1.1” 200 1004
2019-10-24 22:16:23,714 [django.server:124] [basehttp:log_message] [INFO]- “GET /static/assets/js/app.js HTTP/1.1” 200 20844
2019-10-24 22:16:26,509 [django.server:124] [basehttp:log_message] [I

這樣就隻讀取瞭2048個字符,全部讀取的話,循環讀就行

f = open("xxx.log", 'r')
while True:
    block = f.read(2048)
    print(block)
    if not block:
        break
f.close()

2.readline():每次讀取一行,這個方法也不會報錯

f = open("xxx.log", 'r')

while True:
    line = f.readline()
    print(line, end="")
    if not line:
        break
f.close()

3.readlines():讀取全部的行,生成一個list,通過list來對文件進行處理,顯然這種方式依然會造成:MemoyError

真正 Pythonic 的方法

真正 Pythonci 的方法,使用 with 結構打開文件,fp 是一個可迭代對象,可以用 for 遍歷讀取每行的文件內容

with open("xxx.log", 'r') as fp:
    for line in fp:
        print(line, end="")

yield 生成器讀取大文件

前面一篇講yield 生成器的時候提到讀取大文件,函數返回一個可迭代對象,用next()方法讀取文件內容

def read_file(fpath):
    BLOCK_SIZE = 1024
    with open(fpath, 'rb') as f:
        while True:
            block = f.read(BLOCK_SIZE)
            if block:
                yield block
            else:
                return
if __name__ == '__main__':
    a = read_file("xxx.log")
    print(a)            # generator objec
    print(next(a))      # bytes類型
    print(next(a).decode("utf-8"))   # str

運行結果

<generator object read_file at 0x00000226B3005258>
b’\r\n2019-10-24 21:33:31,678 [django.request:93] [base:get_response] [WARNING]- Not Found: /http:/123.125.114.144/\r\n2019-10-24 21:33:31,679 [django.server:124] [basehttp:log_message] [WARNING]- “HEAD http://123.125.114.144/ HTTP/1.1” 404 1678\r\n2019-10-24 22:14:04,121 [django.server:124] [basehttp:log_message] [INFO]- code 400, message Bad request version (\’HTTP\’)\r\n2019-10-24 22:14:04,122 [django.server:124] [basehttp:log_message] [WARNING]- “GET ../../mnt/custom/ProductDefinition HTTP” 400 -\r\n2019-10-24 22:16:21,052 [django.server:124] [basehttp:log_message] [INFO]- “GET /api/login HTTP/1.1” 301 0\r\n2019-10-24 22:16:21,123 [django.server:124] [basehttp:log_message] [INFO]- “GET /api/login/ HTTP/1.1” 200 3876\r\n2019-10-24 22:16:21,192 [django.server:124] [basehttp:log_message] [INFO]- “GET /static/assets/img/main_bg.png HTTP/1.1” 200 2801\r\n2019-10-24 22:16:21,196 [django.server:124] [basehttp:log_message] [INFO]- “GET /static/assets/iconfont/style.css HTTP/1.1” 200 1638\r\n2019-10-24 22:16:21,229 [django.server:124] ‘
[basehttp:log_message] [INFO]- “GET /static/assets/img/bg.jpg HTTP/1.1” 200 135990
2019-10-24 22:16:21,307 [django.server:124] [basehttp:log_message] [INFO]- “GET /static/assets/iconfont/fonts/icomoon.ttf?u4m6fy HTTP/1.1” 200 6900
2019-10-24 22:16:23,525 [django.server:124] [basehttp:log_message] [INFO]- “POST /api/login/ HTTP/1.1” 302 0
2019-10-24 22:16:23,618 [django.server:124] [basehttp:log_message] [INFO]- “GET /api/index/ HTTP/1.1” 200 18447
2019-10-24 22:16:23,709 [django.server:124] [basehttp:log_message] [INFO]- “GET /static/assets/js/commons.js HTTP/1.1” 200 13209
2019-10-24 22:16:23,712 [django.server:124] [basehttp:log_message] [INFO]- “GET /static/assets/css/admin.css HTTP/1.1” 200 19660
2019-10-24 22:16:23,712 [django.server:124] [basehttp:log_message] [INFO]- “GET /static/assets/css/common.css HTTP/1.1” 200 1004
2019-10-24 22:16:23,714 [django.server:124] [basehttp:log_message] [INFO]- “GET /static/assets/js/app.js HTTP/1.1” 200 20844
2019-10-24 22:16:26,509 [django.server:124] [basehtt

到此這篇關於python讀取一個大於10G的txt文件的方法的文章就介紹到這瞭,更多相關python讀取大於10G文件內容請搜索WalkonNet以前的文章或繼續瀏覽下面的相關文章希望大傢以後多多支持WalkonNet!

推薦閱讀: