[InternLM Tutorial] 基础岛第六关 OpenCompass 评测 InternLM-1.8B 实践

运行截图:
在这里插入图片描述
log截图:
在这里插入图片描述
结果截图:
在这里插入图片描述
输出的txt如下:
20240822_141233
tabulate format
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
dataset version metric mode internlm2-chat-1.8b-hf


ceval-computer_network db9ce2 accuracy gen 36.84
ceval-operating_system 1c2571 accuracy gen 42.11
ceval-computer_architecture a74dad accuracy gen 19.05
ceval-college_programming 4ca32a accuracy gen 35.14
ceval-college_physics 963fa8 accuracy gen 31.58
ceval-college_chemistry e78857 accuracy gen 37.50
ceval-advanced_mathematics ce03e2 accuracy gen 31.58
ceval-probability_and_statistics 65e812 accuracy gen 44.44
ceval-discrete_mathematics e894ae accuracy gen 37.50
ceval-electrical_engineer ae42b9 accuracy gen 32.43
ceval-metrology_engineer ee34ea accuracy gen 62.50
ceval-high_school_mathematics 1dc5bf accuracy gen 16.67
ceval-high_school_physics adf25f accuracy gen 36.84
ceval-high_school_chemistry 2ed27f accuracy gen 57.89
ceval-high_school_biology 8e2b9a accuracy gen 26.32
ceval-middle_school_mathematics bee8d5 accuracy gen 26.32
ceval-middle_school_biology 86817c accuracy gen 76.19
ceval-middle_school_physics 8accf6 accuracy gen 57.89
ceval-middle_school_chemistry 167a15 accuracy gen 75.00
ceval-veterinary_medicine b4e08d accuracy gen 60.87
ceval-college_economics f3f4e6 accuracy gen 40.00
ceval-business_administration c1614e accuracy gen 30.30
ceval-marxism cf874c accuracy gen 73.68
ceval-mao_zedong_thought 51c7a4 accuracy gen 66.67
ceval-education_science 591fee accuracy gen 55.17
ceval-teacher_qualification 4e4ced accuracy gen 61.36
ceval-high_school_politics 5c0de2 accuracy gen 52.63
ceval-high_school_geography 865461 accuracy gen 42.11
ceval-middle_school_politics 5be3e7 accuracy gen 80.95
ceval-middle_school_geography 8a63be accuracy gen 75.00
ceval-modern_chinese_history fc01af accuracy gen 56.52
ceval-ideological_and_moral_cultivation a2aa4a accuracy gen 73.68
ceval-logic f5b022 accuracy gen 50.00
ceval-law a110a1 accuracy gen 29.17
ceval-chinese_language_and_literature 0f8b68 accuracy gen 39.13
ceval-art_studies 2a1300 accuracy gen 51.52
ceval-professional_tour_guide 4e673e accuracy gen 62.07
ceval-legal_professional ce8787 accuracy gen 52.17
ceval-high_school_chinese 315705 accuracy gen 42.11
ceval-high_school_history 7eb30a accuracy gen 65.00
ceval-middle_school_history 48ab4a accuracy gen 86.36
ceval-civil_servant 87d061 accuracy gen 44.68
ceval-sports_science 70f27b accuracy gen 47.37
ceval-plant_protection 8941f9 accuracy gen 54.55
ceval-basic_medicine c409d6 accuracy gen 73.68
ceval-clinical_medicine 49e82d accuracy gen 45.45
ceval-urban_and_rural_planner 95b885 accuracy gen 41.30
ceval-accountant 002837 accuracy gen 32.65
ceval-fire_engineer bc23f5 accuracy gen 32.26
ceval-environmental_impact_assessment_engineer c64e2d accuracy gen 45.16
ceval-tax_accountant 3a5e3c accuracy gen 40.82
ceval-physician 6e277d accuracy gen 38.78
ceval-stem - naive_average gen 42.23
ceval-social-science - naive_average gen 57.79
ceval-humanities - naive_average gen 55.25
ceval-other - naive_average gen 45.15
ceval-hard - naive_average gen 36.75
ceval - naive_average gen 48.60
$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$

-------------------------------------------------------------------------------------------------------------------------------- THIS IS A DIVIDER --------------------------------------------------------------------------------------------------------------------------------

csv format
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
dataset,version,metric,mode,internlm2-chat-1.8b-hf
ceval-computer_network,db9ce2,accuracy,gen,36.84
ceval-operating_system,1c2571,accuracy,gen,42.11
ceval-computer_architecture,a74dad,accuracy,gen,19.05
ceval-college_programming,4ca32a,accuracy,gen,35.14
ceval-college_physics,963fa8,accuracy,gen,31.58
ceval-college_chemistry,e78857,accuracy,gen,37.50
ceval-advanced_mathematics,ce03e2,accuracy,gen,31.58
ceval-probability_and_statistics,65e812,accuracy,gen,44.44
ceval-discrete_mathematics,e894ae,accuracy,gen,37.50
ceval-electrical_engineer,ae42b9,accuracy,gen,32.43
ceval-metrology_engineer,ee34ea,accuracy,gen,62.50
ceval-high_school_mathematics,1dc5bf,accuracy,gen,16.67
ceval-high_school_physics,adf25f,accuracy,gen,36.84
ceval-high_school_chemistry,2ed27f,accuracy,gen,57.89
ceval-high_school_biology,8e2b9a,accuracy,gen,26.32
ceval-middle_school_mathematics,bee8d5,accuracy,gen,26.32
ceval-middle_school_biology,86817c,accuracy,gen,76.19
ceval-middle_school_physics,8accf6,accuracy,gen,57.89
ceval-middle_school_chemistry,167a15,accuracy,gen,75.00
ceval-veterinary_medicine,b4e08d,accuracy,gen,60.87
ceval-college_economics,f3f4e6,accuracy,gen,40.00
ceval-business_administration,c1614e,accuracy,gen,30.30
ceval-marxism,cf874c,accuracy,gen,73.68
ceval-mao_zedong_thought,51c7a4,accuracy,gen,66.67
ceval-education_science,591fee,accuracy,gen,55.17
ceval-teacher_qualification,4e4ced,accuracy,gen,61.36
ceval-high_school_politics,5c0de2,accuracy,gen,52.63
ceval-high_school_geography,865461,accuracy,gen,42.11
ceval-middle_school_politics,5be3e7,accuracy,gen,80.95
ceval-middle_school_geography,8a63be,accuracy,gen,75.00
ceval-modern_chinese_history,fc01af,accuracy,gen,56.52
ceval-ideological_and_moral_cultivation,a2aa4a,accuracy,gen,73.68
ceval-logic,f5b022,accuracy,gen,50.00
ceval-law,a110a1,accuracy,gen,29.17
ceval-chinese_language_and_literature,0f8b68,accuracy,gen,39.13
ceval-art_studies,2a1300,accuracy,gen,51.52
ceval-professional_tour_guide,4e673e,accuracy,gen,62.07
ceval-legal_professional,ce8787,accuracy,gen,52.17
ceval-high_school_chinese,315705,accuracy,gen,42.11
ceval-high_school_history,7eb30a,accuracy,gen,65.00
ceval-middle_school_history,48ab4a,accuracy,gen,86.36
ceval-civil_servant,87d061,accuracy,gen,44.68
ceval-sports_science,70f27b,accuracy,gen,47.37
ceval-plant_protection,8941f9,accuracy,gen,54.55
ceval-basic_medicine,c409d6,accuracy,gen,73.68
ceval-clinical_medicine,49e82d,accuracy,gen,45.45
ceval-urban_and_rural_planner,95b885,accuracy,gen,41.30
ceval-accountant,002837,accuracy,gen,32.65
ceval-fire_engineer,bc23f5,accuracy,gen,32.26
ceval-environmental_impact_assessment_engineer,c64e2d,accuracy,gen,45.16
ceval-tax_accountant,3a5e3c,accuracy,gen,40.82
ceval-physician,6e277d,accuracy,gen,38.78
ceval-stem,-,naive_average,gen,42.23
ceval-social-science,-,naive_average,gen,57.79
ceval-humanities,-,naive_average,gen,55.25
ceval-other,-,naive_average,gen,45.15
ceval-hard,-,naive_average,gen,36.75
ceval,-,naive_average,gen,48.60
$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$

-------------------------------------------------------------------------------------------------------------------------------- THIS IS A DIVIDER --------------------------------------------------------------------------------------------------------------------------------

raw format
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^

Model: internlm2-chat-1.8b-hf
ceval-computer_network: {‘accuracy’: 36.84210526315789}
ceval-operating_system: {‘accuracy’: 42.10526315789473}
ceval-computer_architecture: {‘accuracy’: 19.047619047619047}
ceval-college_programming: {‘accuracy’: 35.13513513513514}
ceval-college_physics: {‘accuracy’: 31.57894736842105}
ceval-college_chemistry: {‘accuracy’: 37.5}
ceval-advanced_mathematics: {‘accuracy’: 31.57894736842105}
ceval-probability_and_statistics: {‘accuracy’: 44.44444444444444}
ceval-discrete_mathematics: {‘accuracy’: 37.5}
ceval-electrical_engineer: {‘accuracy’: 32.432432432432435}
ceval-metrology_engineer: {‘accuracy’: 62.5}
ceval-high_school_mathematics: {‘accuracy’: 16.666666666666664}
ceval-high_school_physics: {‘accuracy’: 36.84210526315789}
ceval-high_school_chemistry: {‘accuracy’: 57.89473684210527}
ceval-high_school_biology: {‘accuracy’: 26.31578947368421}
ceval-middle_school_mathematics: {‘accuracy’: 26.31578947368421}
ceval-middle_school_biology: {‘accuracy’: 76.19047619047619}
ceval-middle_school_physics: {‘accuracy’: 57.89473684210527}
ceval-middle_school_chemistry: {‘accuracy’: 75.0}
ceval-veterinary_medicine: {‘accuracy’: 60.86956521739131}
ceval-college_economics: {‘accuracy’: 40.0}
ceval-business_administration: {‘accuracy’: 30.303030303030305}
ceval-marxism: {‘accuracy’: 73.68421052631578}
ceval-mao_zedong_thought: {‘accuracy’: 66.66666666666666}
ceval-education_science: {‘accuracy’: 55.172413793103445}
ceval-teacher_qualification: {‘accuracy’: 61.36363636363637}
ceval-high_school_politics: {‘accuracy’: 52.63157894736842}
ceval-high_school_geography: {‘accuracy’: 42.10526315789473}
ceval-middle_school_politics: {‘accuracy’: 80.95238095238095}
ceval-middle_school_geography: {‘accuracy’: 75.0}
ceval-modern_chinese_history: {‘accuracy’: 56.52173913043478}
ceval-ideological_and_moral_cultivation: {‘accuracy’: 73.68421052631578}
ceval-logic: {‘accuracy’: 50.0}
ceval-law: {‘accuracy’: 29.166666666666668}
ceval-chinese_language_and_literature: {‘accuracy’: 39.130434782608695}
ceval-art_studies: {‘accuracy’: 51.515151515151516}
ceval-professional_tour_guide: {‘accuracy’: 62.06896551724138}
ceval-legal_professional: {‘accuracy’: 52.17391304347826}
ceval-high_school_chinese: {‘accuracy’: 42.10526315789473}
ceval-high_school_history: {‘accuracy’: 65.0}
ceval-middle_school_history: {‘accuracy’: 86.36363636363636}
ceval-civil_servant: {‘accuracy’: 44.680851063829785}
ceval-sports_science: {‘accuracy’: 47.368421052631575}
ceval-plant_protection: {‘accuracy’: 54.54545454545454}
ceval-basic_medicine: {‘accuracy’: 73.68421052631578}
ceval-clinical_medicine: {‘accuracy’: 45.45454545454545}
ceval-urban_and_rural_planner: {‘accuracy’: 41.30434782608695}
ceval-accountant: {‘accuracy’: 32.6530612244898}
ceval-fire_engineer: {‘accuracy’: 32.25806451612903}
ceval-environmental_impact_assessment_engineer: {‘accuracy’: 45.16129032258064}
ceval-tax_accountant: {‘accuracy’: 40.816326530612244}
ceval-physician: {‘accuracy’: 38.775510204081634}
ceval-stem: {‘naive_average’: 42.23273800933983}
ceval-social-science: {‘naive_average’: 57.78791807103967}
ceval-humanities: {‘naive_average’: 55.24818006394801}
ceval-other: {‘naive_average’: 45.1547348424325}
ceval-hard: {‘naive_average’: 36.75073099415205}
ceval: {‘naive_average’: 48.59550009360344}
$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$

评论
添加红包

请填写红包祝福语或标题

红包个数最小为10个

红包金额最低5元

当前余额3.43前往充值 >
需支付:10.00
成就一亿技术人!
领取后你会自动成为博主和红包主的粉丝 规则
hope_wisdom
发出的红包
实付
使用余额支付
点击重新获取
扫码支付
钱包余额 0

抵扣说明:

1.余额是钱包充值的虚拟货币,按照1:1的比例进行支付金额的抵扣。
2.余额无法直接购买下载,可以购买VIP、付费专栏及课程。

余额充值