多个coco数据标注文件合并

小张Tt

已于 2024-03-22 11:07:20 修改

阅读量1.9k

点赞数 14

分类专栏：数据处理与图像处理文章标签：人工智能计算机视觉深度学习

于 2024-01-22 17:33:13 首次发布

本文链接：https://blog.csdn.net/weixin_43788282/article/details/135753667

版权

数据处理与图像处理专栏收录该内容

22 篇文章

订阅专栏

本文介绍了一个Python脚本，用于合并COCO格式的JSON文件，以方便管理目标检测和图像分割任务中的数据。脚本通过更新图像和注释ID，将多个文件整合到一个merged_coco.json文件中。

摘要生成于 C知道，由 DeepSeek-R1 满血版支持，前往体验 >

一、coco数据集是什么？

COCO（Common Objects in Context）是一个用于目标检测和图像分割任务的标注格式。如果你有多个COCO格式的JSON文件，你可能需要将它们合并成一个文件，以便更方便地处理和管理数据。在这篇博客中，我们将介绍一个用Python编写的脚本，可以实现这一合并操作。

二、完整代码

import json
import os

def merge_coco_files(folder_path):
    merged_data = {
        "info": {
            "year": 2023,
            "version": "1",
            "date_created": "no need record"
        },
        "images": [],
        "annotations": [],
        "licenses": [
            {
                "id": 1,
                "name": "Unknown",
                "url": ""
            }
        ],
        "categories": [
            {
                "id": 1,
                "name": "hd",
                "supercategory": ""
            }
        ]
    }

    image_id_counter = 1
    annotation_id_counter = 1

    for root, dirs, files in os.walk(folder_path):
        for file in files:
            if file.endswith(".json"):
                file_path = os.path.join(root, file)
                with open(file_path, 'r') as f:
                    data = json.load(f)

                    # Update image IDs and filenames
                    for image in data["images"]:
                        image["id"] = image_id_counter
                        image_id_counter += 1

                        # Use the original file name from the COCO file
                        image["file_name"] = image["file_name"]

                        # Append the updated image to the merged_data only if it's not already present
                        if image not in merged_data["images"]:
                            merged_data["images"].append(image)

                    # Update annotation IDs and image IDs
                    for annotation in data["annotations"]:
                        annotation["id"] = annotation_id_counter
                        annotation_id_counter += 1
                        annotation["image_id"] = image_id_counter - 1  # Use the last assigned image ID

                        # Append the updated annotation to the merged_data
                        merged_data["annotations"].append(annotation)

    # Save the merged data to a new JSON file
    output_path = os.path.join(folder_path, "merged_coco.json")
    with open(output_path, 'w') as output_file:
        json.dump(merged_data, output_file, indent=4)

    print(f'Merged data saved to: {output_path}')

# Provide the path to the folder containing the COCO JSON files
folder_path = r''
merge_coco_files(folder_path)