進捗報告資料原因分類モデルの baseline 構築

佐藤 真 2026年5月28日(木)

01前回まで

とりあえずモデルを作成し、タスク自体の難易度を測る

02実験設定

タスク

10 カテゴリ

# 正式名称 データファイル内の表記
1LexicalLexical
2ImplicatureImplicature
3PresuppositionPresupposition
4Probabilistic EnrichmentProbabilistic
5ImperfectionImperfection
6CoreferenceCoreference
7Temporal ReferenceTemporal
8Interrogative HypothesisInterrogative.
9Accommodating Minimally Added ContentAccommodating
10High OverlapHigh Overlap

データ分割

taxonomy_round1_release.jsonl(400 件)→ train: 400 件(N/A は空ラベルとして使用) taxonomy_round2_release.jsonl(110 件)→ 前半 60 件(MNLI 由来)→ test 後半 50 件(ChaosNLI 由来)→ valid

使用モデル

本実験では BERT-base-uncasedDeBERTa-v3-base の 2 モデルを比較する。

モデル 理由
BERT-base-uncasedTransformer 系 fine-tune の標準モデル
DeBERTa-v3-baseBERT の後継

共通設定

パラメータ
MAX_LENGTH256
BATCH_SIZE16
EPOCHS15
THRESHOLD0.5
オプティマイザAdamW (weight_decay=0.01)
スケジューラLinear warmup (warmup ratio=10%)
Gradient clippingmax_norm=1.0
損失関数BCEWithLogitsLoss(pos_weight = 負例数/正例数、ラベル別)
SEED42

03実験結果

valid セット(ChaosNLI, n = 50)

指標 BERT fine-tune DeBERTa fine-tune
Micro F10.39020.3605
Macro F10.12280.1391
Exact Match0.16000.1200

ラベル別 F1(valid)

カテゴリ BERT DeBERTa support
Lexical0.37840.540518
Implicature0.00000.00001
Presupposition0.00000.00001
Probabilistic Enrichment0.62750.533320
Imperfection0.00000.00002
Coreference0.00000.250011
Temporal Reference0.00000.00003
Interrogative Hypothesis0.00000.00000
Accommodating Minimally Added Content0.22220.06672
High Overlap0.00000.00001

test セット(MNLI, n = 60)

指標 BERT fine-tune DeBERTa fine-tune
Micro F10.32860.3959
Macro F10.19370.2797
Exact Match0.15000.0667

ラベル別 F1(test)

カテゴリ BERT DeBERTa support
Lexical0.24390.390217
Implicature0.00000.00000
Presupposition0.00000.25003
Probabilistic Enrichment0.47060.566719
Imperfection0.00000.28574
Coreference0.00000.333310
Temporal Reference0.00000.00003
Interrogative Hypothesis1.00000.83335
Accommodating Minimally Added Content0.22220.13793
High Overlap0.00000.00001