進捗報告資料原因分類モデルの baseline 構築
01前回まで
とりあえずモデルを作成し、タスク自体の難易度を測る
- train=400 / valid=50 として、test に関しては今回は使わずに検証のみ精度を見る
- タスクとしての難易度を図るため、シンプルなよく使われるモデルを利用する
02実験設定
タスク
- 入力:(premise, hypothesis) テキストペア
- 出力:不一致原因の 10 カテゴリのマルチラベル予測
- 評価指標:Micro F1 / Macro F1 / Exact Match / ラベル別 F1
10 カテゴリ
| # |
正式名称 |
データファイル内の表記 |
| 1 | Lexical | Lexical |
| 2 | Implicature | Implicature |
| 3 | Presupposition | Presupposition |
| 4 | Probabilistic Enrichment | Probabilistic |
| 5 | Imperfection | Imperfection |
| 6 | Coreference | Coreference |
| 7 | Temporal Reference | Temporal |
| 8 | Interrogative Hypothesis | Interrogative. |
| 9 | Accommodating Minimally Added Content | Accommodating |
| 10 | High Overlap | High Overlap |
データ分割
taxonomy_round1_release.jsonl(400 件)→ train: 400 件(N/A は空ラベルとして使用)
taxonomy_round2_release.jsonl(110 件)→ 前半 60 件(MNLI 由来)→ test
後半 50 件(ChaosNLI 由来)→ valid
使用モデル
本実験では BERT-base-uncased と DeBERTa-v3-base の 2 モデルを比較する。
| モデル |
理由 |
| BERT-base-uncased | Transformer 系 fine-tune の標準モデル |
| DeBERTa-v3-base | BERT の後継 |
共通設定
| パラメータ |
値 |
| MAX_LENGTH | 256 |
| BATCH_SIZE | 16 |
| EPOCHS | 15 |
| THRESHOLD | 0.5 |
| オプティマイザ | AdamW (weight_decay=0.01) |
| スケジューラ | Linear warmup (warmup ratio=10%) |
| Gradient clipping | max_norm=1.0 |
| 損失関数 | BCEWithLogitsLoss(pos_weight = 負例数/正例数、ラベル別) |
| SEED | 42 |
03実験結果
valid セット(ChaosNLI, n = 50)
| 指標 |
BERT fine-tune |
DeBERTa fine-tune |
| Micro F1 | 0.3902 | 0.3605 |
| Macro F1 | 0.1228 | 0.1391 |
| Exact Match | 0.1600 | 0.1200 |
ラベル別 F1(valid)
| カテゴリ |
BERT |
DeBERTa |
support |
| Lexical | 0.3784 | 0.5405 | 18 |
| Implicature | 0.0000 | 0.0000 | 1 |
| Presupposition | 0.0000 | 0.0000 | 1 |
| Probabilistic Enrichment | 0.6275 | 0.5333 | 20 |
| Imperfection | 0.0000 | 0.0000 | 2 |
| Coreference | 0.0000 | 0.2500 | 11 |
| Temporal Reference | 0.0000 | 0.0000 | 3 |
| Interrogative Hypothesis | 0.0000 | 0.0000 | 0 |
| Accommodating Minimally Added Content | 0.2222 | 0.0667 | 2 |
| High Overlap | 0.0000 | 0.0000 | 1 |
test セット(MNLI, n = 60)
| 指標 |
BERT fine-tune |
DeBERTa fine-tune |
| Micro F1 | 0.3286 | 0.3959 |
| Macro F1 | 0.1937 | 0.2797 |
| Exact Match | 0.1500 | 0.0667 |
ラベル別 F1(test)
| カテゴリ |
BERT |
DeBERTa |
support |
| Lexical | 0.2439 | 0.3902 | 17 |
| Implicature | 0.0000 | 0.0000 | 0 |
| Presupposition | 0.0000 | 0.2500 | 3 |
| Probabilistic Enrichment | 0.4706 | 0.5667 | 19 |
| Imperfection | 0.0000 | 0.2857 | 4 |
| Coreference | 0.0000 | 0.3333 | 10 |
| Temporal Reference | 0.0000 | 0.0000 | 3 |
| Interrogative Hypothesis | 1.0000 | 0.8333 | 5 |
| Accommodating Minimally Added Content | 0.2222 | 0.1379 | 3 |
| High Overlap | 0.0000 | 0.0000 | 1 |