有关OCR的数据集整理

最新推荐文章于 2025-03-12 19:24:11 发布

森森之火

最新推荐文章于 2025-03-12 19:24:11 发布

阅读量711

点赞数

分类专栏：人工智能文章标签： ocr

原文链接：https://www.mianshigee.com/project/WenmuZhou-OCR_DataSet

版权

人工智能专栏收录该内容

8 篇文章

订阅专栏

提供数据集百度云链接
数据集转换为统一格式(检测和识别)
- icdar2015
- MLT2019
- COCO-Text_v2
- ReCTS
- SROIE
- ArT
- LSVT
- Synth800k
- icdar2017rctw
- MTWI 2018
- 百度中文场景文字识别
- mjsynth
- Synthetic Chinese String Dataset(360万中文数据集)
提供读取脚本

参考收集并整理有关OCR的数据集，以便制作通用OCR-面圈网 (mianshigee.com)

下载

百度云提取码：9s4x

数据集

数据集	主页	适用情况	数据情况	标注形式	说明
ICDAR2015	Overview - Incidental Scene Text - Robust Reading Competition	检测&识别	语言: 英文 train:1,000 test:500	x1, y1, x2, y2, x3, y3, x4, y4, transcription	坐标: x1, y1, x2, y2, x3, y3, x4, y4 transcription : 框内的文字信息
MLT2019	Overview - ICDAR 2019 Robust Reading Challenge on Multi-lingual scene text detection and recognition - Robust Reading Competition	检测&识别	语言: 混合 train:10,000 test:10,000	x1,y1,x2,y2,x3,y3,x4,y4,script,transcription	坐标: x1, y1, x2, y2, x3, y3, x4, y4 script: 文字所属语言 transcription : 框内的文字信息
COCO-Text_v2	COCO-Text V2.0	检测&识别	语言: 混合 train:43,686 validation:10,000 test:10,000	json
ReCTS	Overview - ICDAR 2019 Robust Reading Challenge on Reading Chinese Text on Signboard - Robust Reading Competition	检测&识别	语言: 混合 train:20,000 test:5,000	{ “chars”: [ {“points”: [x1,y1,x2,y2,x3,y3,x4,y4], “transcription” : “trans1”, "ignore":0 }, {“points”: [x1,y1,x2,y2,x3,y3,x4,y4], “transcription” : “trans2”, " ignore ":0 }], “lines”: [ {“points”: [x1,y1,x2,y2,x3,y3,x4,y4] , “transcription” : “trans3”, "ignore ":0 }], }	points: x1,y1,x2,y2,x3,y3,x4,y4 chars: 字符级别的标注 lines: 行级别的标注. transcription : 框内的文字信息 ignore: 0:不忽略，1:忽略
SROIE	Overview - ICDAR 2019 Robust Reading Challenge on Scanned Receipts OCR and Information Extraction - Robust Reading Competition	检测&识别	语言: 英文 train:699 test:400	x1, y1, x2, y2, x3, y3, x4, y4, transcription	坐标: x1, y1, x2, y2, x3, y3, x4, y4 transcription : 框内的文字信息
ArT(已包含Total-Text和SCUT-CTW1500)	Overview - ICDAR2019 Robust Reading Challenge on Arbitrary-Shaped Text - Robust Reading Competition	检测&识别	语言: 混合 train: 5,603 test: 4,563	{ “gt_1”: [ {“points”: [[x1, y1], [x2, y2], …, [xn, yn]], “transcription” : “trans1”, “language” : “Latin”, "illegibility": false }, {“points”: [[x1, y1], [x2, y2], …, [xn, yn]], “transcription” : “trans2”, “language” : “Chinese”, "illegibility": false }], }	points: x1,y1,x2,y2,x3,y3,x4,y4…xn,yn transcription : 框内的文字信息 language: 语言信息 illegibility: 是否模糊
LSVT	Overview - ICDAR2019 Robust Reading Challenge on Large-scale Street View Text with Partial Labeling - Robust Reading Competition	检测&识别	语言: 混合全标注 train: 30,000 test: 20,000 只标注文本 400,000	{ “gt_1”: [ {“points”: [[x1, y1], [x2, y2], …, [xn, yn]], “transcription” : “trans1”, "illegibility": false }, {“points”: [[x1, y1], [x2, y2], …, [xn, yn]], “transcription” : “trans2”, "illegibility": false }], }	points: x1,y1,x2,y2,x3,y3,x4,y4…xn,yn transcription : 框内的文字信息 illegibility: 是否模糊
Synth800k	Visual Geometry Group - University of Oxford	检测&识别	语言: 英文 800,000	imnames: wordBB: charBB: txt:	imnames: 文件名称 wordBB: 24n,每张图像内的文本框 charBB: 24n,每张图像内的字符框 txt: 每张图形内的字符串
icdar2017rctw	ICDAR 2017 RCTW 中文场景文本检测和识别数据集_icdar2017_忘泪的博客-CSDN博客	检测&识别	语言: 混合 train:8,034 test:4,229	x1,y1,x2,y2,x3,y3,x4,y4,<识别难易程度>,transcription	坐标: x1, y1, x2, y2, x3, y3, x4, y4 transcription : 框内的文字信息
MTWI 2018	识别: https://tianchi.aliyun.com/competition/entrance/231684/introduction 检测: https://tianchi.aliyun.com/competition/entrance/231685/introduction	检测&识别	语言: 混合 train:10,000 test:10,000	x1, y1, x2, y2, x3, y3, x4, y4, transcription	坐标: x1, y1, x2, y2, x3, y3, x4, y4 transcription : 框内的文字信息
百度中文场景文字识别	飞桨学习赛：中文场景文字识别（本赛事已结束，可前往同名比赛进行学习） - 飞桨AI Studio	识别	语言: 混合 train:未统计 test:未统计	h,w,name,value	h: 图片高度 w: 图片宽度 name: 图片名 value: 图片上文字
mjsynth	Visual Geometry Group - University of Oxford	识别	语言: 英文 9,000,000	-	-
Synthetic Chinese String Dataset(360万中文数据集)	链接：百度网盘-链接不存在提取码：spyi	识别	语言: 混合 300k	-	-
英文识别数据大礼包(GitHub - clovaai/deep-text-recognition-benchmark: Text recognition (optical character recognition) with deep learning methods.) 训练：MJSynth和SynthText 验证：IIIT, SVT, IC03, IC13, IC15, SVTP, CUTE	链接：百度网盘请输入提取码提取码：rryk	识别	语言: 英文	-	-

数据生成工具

https://github.com/TianzhongSong/awesome-SynthText

评论

被折叠的条评论为什么被折叠?

到【灌水乐园】发言

查看更多评论

添加红包

成就一亿技术人!

hope_wisdom

发出的红包

实付元

使用余额支付

点击重新获取

扫码支付

钱包余额 0

抵扣说明：

1.余额是钱包充值的虚拟货币，按照1:1的比例进行支付金额的抵扣。
2.余额无法直接购买下载，可以购买VIP、付费专栏及课程。