モデル評価をClaude Code x Agent Skillsを使って完全自動運転した話
こんにちは!逆瀬川 (@gyakuse) です!
今日は機械学習モデルの評価を自動で行う仕組みをAgent Skillsで作ってみたので共有しようと思います。これをあくまで叩き台として改善していったら、めっちゃ便利になる気がします。
これはどんなものか?
Claude Codeセッションを立ち上げて、/agentic-bench {URLやモデル名など} で実行するだけで、かんたんなレポートが出てきます。忙しいときに便利!
レポートの構造や評価アプローチを詳細にしたりするときは、プロンプトで指示してもいいですし、今回作ったものをforkして改善しても良いと思います。Claude Codeのみ検証していてディレクトリ構造も ./claude/skills 依存ですが、symlinkして普通に動くと思います。
今回作ったSkillはMITライセンスで公開しています。
レポートはGitHub Pagesで公開しています。
こんな感じ:




使い方
- https://github.com/nyosegawa/agentic-bench をclone
-
.env.exampleをコピーして.envを作成し、各種API KEYを入力 - Claudeなどのセッションを開始
-
/agentic-bench {URLやモデル名など}で実行- 例:
/agentic-bench https://huggingface.co/zai-org/GLM-OCR - 日本語指示入れたら、レポートも日本語出力になると思います
- 例:
- 予算を承認
- レポートを見る
対応しているGPUサーバー
基本的にクラウドGPUサービスを使って推論を実行します。.env に設定したトークンの有無を自動検出し、コスト最安のプロバイダを選択します。
| プロバイダ | 料金 | 特徴 |
|---|---|---|
| HF Inference API | 無料 (HF Pro $9/月) | HuggingFace上の対応モデルをそのまま叩ける。最もシンプル |
| HF Inference Endpoints | $0.50–2.50/hr | 任意のHFモデルを専用GPUにデプロイ。HF_TOKENだけで使える |
| Modal | $0.59–3.95/hr | サーバーレスGPU。月$30の無料クレジットあり |
| beam.cloud | $0.54–3.50/hr | 専用エンドポイント方式。既存クレジットで従量課金 |
| Vast.ai | $0.10–2.00/hr | 分散型GPUマーケットプレイス。最安値でGPUを借りられる |
| RunPod | $0.34–2.69/hr | Pod(永続VM)方式。秒単位の従量課金でバランスが良い |
最小構成は HF_TOKEN だけでOKです。HF Inference API と HF Inference Endpoints が使えます。
実際の例
呼び出し
/agentic-bench https://huggingface.co/zai-org/GLM-OCR
計画と予算の確認
Phase 1: Research Summary — GLM-OCR
┌────────────────┬──────────────────────────────────────────────────────────────────┐
│ 項目 │ 詳細 │
├────────────────┼──────────────────────────────────────────────────────────────────┤
│ モデル │ zai-org/GLM-OCR │
├────────────────┼──────────────────────────────────────────────────────────────────┤
│ パラメータ数 │ 0.9B │
├────────────────┼──────────────────────────────────────────────────────────────────┤
│ アーキテクチャ │ GLM-V encoder-decoder (CogViT visual encoder + GLM-0.5B decoder) │
├────────────────┼──────────────────────────────────────────────────────────────────┤
│ タスク │ ドキュメントOCR(テキスト/数式/テーブル認識 + 情報抽出) │
├────────────────┼──────────────────────────────────────────────────────────────────┤
│ ライセンス │ MIT │
├────────────────┼──────────────────────────────────────────────────────────────────┤
│ 対応言語 │ 中/英/仏/西/露/独/日/韓 │
├────────────────┼──────────────────────────────────────────────────────────────────┤
│ ベンチマーク │ OmniDocBench V1.5: 94.62 (1位) │
├────────────────┼──────────────────────────────────────────────────────────────────┤
│ 推定VRAM │ ~2GB (fp16) — T4 (16GB) で十分 │
└────────────────┴──────────────────────────────────────────────────────────────────┘
評価計画
GLM-OCRはOCR特化モデルなので、VLM汎用テストではなくOCR固有のテストを実施:
1. Smoke test: モデルロード + 簡単なテキスト画像認識
2. テキスト認識: 英語・日本語のドキュメント画像
3. 数式認識: LaTeX数式を含む画像
4. テーブル認識: 表組みの画像
5. 情報抽出: 構造化データ抽出(JSON指定)
6. パフォーマンス測定: 5回実行のメディアン
コスト見積もり
┌────────────┬─────────┬────────────┬───────────────────────────────┐
│ プロバイダ │ GPU │ 推定コスト │ 備考 │
├────────────┼─────────┼────────────┼───────────────────────────────┤
│ Modal (T4) │ T4 16GB │ ~$0.15 │ 月$30無料クレジット内 │
├────────────┼─────────┼────────────┼───────────────────────────────┤
│ Modal (L4) │ L4 24GB │ ~$0.20 │ 月$30無料クレジット内 │
├────────────┼─────────┼────────────┼───────────────────────────────┤
│ Colab (T4) │ T4 16GB │ ~$0.29 │ Pro含有コンピュートユニット内 │
└────────────┴─────────┴────────────┴───────────────────────────────┘
推奨: Modal T4 (最安、無料クレジット内で実行可能)
実行に進みますか?
レポート
⏺ GLM-OCR ベンチマーク完了
結果サマリー
┌─────────────────┬───────┬──────────────────────────────────┐
│ テスト │ 結果 │ 詳細 │
├─────────────────┼───────┼──────────────────────────────────┤
│ Smoke Test │ PASS │ 16.95s でロード、VRAM 2.34GB │
├─────────────────┼───────┼──────────────────────────────────┤
│ English Text │ PASS │ 3行の英文を完全に認識 │
├─────────────────┼───────┼──────────────────────────────────┤
│ Japanese Text │ PASS │ 日本語テキスト完全認識 │
├─────────────────┼───────┼──────────────────────────────────┤
│ Math Formula │ PASS* │ 正しいLaTeX出力(E = m c ^ {2}) │
├─────────────────┼───────┼──────────────────────────────────┤
│ Table │ PASS │ HTML形式の構造化テーブル出力 │
├─────────────────┼───────┼──────────────────────────────────┤
│ Info Extraction │ PASS │ JSON形式で全フィールド正確抽出 │
└─────────────────┴───────┴──────────────────────────────────┘
パフォーマンス (T4 GPU)
- スループット: 4.1 tok/s (median, raw transformers)
- レイテンシ: 3.138s (p50) / 4.14s (p99)
- VRAM使用量: 2.34 GB / 15.6 GB (15%)
- 推定コスト: ~$0.15 (Modal無料クレジット内)
所感
- 0.9Bパラメータで驚異的な精度 — テキスト・テーブル・情報抽出すべて完璧
- 日本語対応も良好 — 完全な認識結果
- 非常に軽量 — T4どころかコンシューマGPUでも動作可能
- 注意点: transformers のdev版が必要(5.2.0にはバグあり)、vLLM/SGLang推奨
レポートはブラウザで開きました: results/2026-02-21_glm-ocr/report.html
スキルを作ったり改善したりするときの話
作ったスキルを改善する
Agentの能力の改善はTranscript(つまり、実行のログ)を評価するのが一番です。 Manusの研究者も言っていましたが、動いている様子をつぶさに見て、判断するヒューリスティックな方法も正しいと思います。今回は直接呼ぶようなSkillなので、ルーティング問題を回避できます。わたしの場合は単純に実行したログをそのままコピペして、エッジケース対応にならないように修正してと伝えました。だいたいこの指示をしてもエッジケース対応してくるので(許さん)、調整します。
Transcript評価の罠
ただし、Transcript評価は後知恵バイアスが働きます。最適なルートのための改善を考えてしまいがちです。重要なのは、探索的に試行錯誤をしながら結果を出すようなAgenticなはたらきを阻害しないようなSkillを作ることです。
自明であるような処理が本当に自明かはよく考えなければなりませんし、Agentの自由度を無駄に阻害するような指示・scriptはゴミです。気をつけてやる必要があります。
Appendix: 実行ログ
見てて楽しいので共有します。
~/src/github.com/nyosegawa/agentic-bench main [?]
❯ claude 15:00
✻
|
▟█▙ Claude Code v2.1.50
▐▛███▜▌ Opus 4.6 (1M context) · Claude Max
▝▜█████▛▘ ~/src/github.com/nyosegawa/agentic-bench
▘▘ ▝▝
❯ /agentic-bench https://huggingface.co/zai-org/GLM-OCR
⏺ GLM-OCR モデルの検証を開始します。まず Phase 1: Research からスタートします。
⏺ Searched for 1 pattern, read 4 files (ctrl+o to expand)
⏺ Fetch(https://huggingface.co/zai-org/GLM-OCR)
⎿ Received 141.3KB (200 OK)
⏺ Searched for 1 pattern, read 1 file (ctrl+o to expand)
⏺ ---
Phase 1: Research Summary — GLM-OCR
┌────────────────┬──────────────────────────────────────────────────────────────────┐
│ 項目 │ 詳細 │
├────────────────┼──────────────────────────────────────────────────────────────────┤
│ モデル │ zai-org/GLM-OCR │
├────────────────┼──────────────────────────────────────────────────────────────────┤
│ パラメータ数 │ 0.9B │
├────────────────┼──────────────────────────────────────────────────────────────────┤
│ アーキテクチャ │ GLM-V encoder-decoder (CogViT visual encoder + GLM-0.5B decoder) │
├────────────────┼──────────────────────────────────────────────────────────────────┤
│ タスク │ ドキュメントOCR(テキスト/数式/テーブル認識 + 情報抽出) │
├────────────────┼──────────────────────────────────────────────────────────────────┤
│ ライセンス │ MIT │
├────────────────┼──────────────────────────────────────────────────────────────────┤
│ 対応言語 │ 中/英/仏/西/露/独/日/韓 │
├────────────────┼──────────────────────────────────────────────────────────────────┤
│ ベンチマーク │ OmniDocBench V1.5: 94.62 (1位) │
├────────────────┼──────────────────────────────────────────────────────────────────┤
│ 推定VRAM │ ~2GB (fp16) — T4 (16GB) で十分 │
└────────────────┴──────────────────────────────────────────────────────────────────┘
評価計画
GLM-OCRはOCR特化モデルなので、VLM汎用テストではなくOCR固有のテストを実施:
1. Smoke test: モデルロード + 簡単なテキスト画像認識
2. テキスト認識: 英語・日本語のドキュメント画像
3. 数式認識: LaTeX数式を含む画像
4. テーブル認識: 表組みの画像
5. 情報抽出: 構造化データ抽出(JSON指定)
6. パフォーマンス測定: 5回実行のメディアン
コスト見積もり
┌────────────┬─────────┬────────────┬───────────────────────────────┐
│ プロバイダ │ GPU │ 推定コスト │ 備考 │
├────────────┼─────────┼────────────┼───────────────────────────────┤
│ Modal (T4) │ T4 16GB │ ~$0.15 │ 月$30無料クレジット内 │
├────────────┼─────────┼────────────┼───────────────────────────────┤
│ Modal (L4) │ L4 24GB │ ~$0.20 │ 月$30無料クレジット内 │
├────────────┼─────────┼────────────┼───────────────────────────────┤
│ Colab (T4) │ T4 16GB │ ~$0.29 │ Pro含有コンピュートユニット内 │
└────────────┴─────────┴────────────┴───────────────────────────────┘
推奨: Modal T4 (最安、無料クレジット内で実行可能)
実行に進みますか?
⏺ User answered Claude's questions:
⎿ · Phase 2 (GPU実行) に進みますか?Modal T4で推定$0.15です。 → 実行する (Recommended)
⏺ Phase 2: 実行フェーズに進みます。まずModal参照ガイドを確認し、推論スクリプトを準備します。
⏺ Read 1 file (ctrl+o to expand)
⏺ Bash(mkdir -p /Users/sakasegawa/src/github.com/nyosegawa/agentic-bench/results/2026-02-21_glm-ocr/{workspace,artifacts})
⎿ Done
⏺ テスト画像の準備と推論スクリプトを作成します。OCRモデルなので、テキスト・数式・テーブル・日本語を含む多様なテスト画像が必要です。
⏺ Write(results/2026-02-21_glm-ocr/workspace/run.py)
⎿ Wrote 413 lines to results/2026-02-21_glm-ocr/workspace/run.py
1 """GLM-OCR Benchmark on Modal (T4 GPU)
2
3 Tests:
4 1. Smoke test: Simple text recognition
5 2. English text: Document paragraph
6 3. Japanese text: Japanese document
7 4. Math formula: LaTeX formula image
8 5. Table: Structured table image
9 6. Information extraction: Structured JSON extraction
10 7. Performance: 5-run speed measurement
… +403 lines (ctrl+o to expand)
⏺ Bash(modal run results/2026-02-21_glm-ocr/workspace/run.py 2>&1)
⎿ Error: Exit code 1
╭─ Error ──────────────────────────────────────────────────────────────────────╮
│ Token missing. Could not authenticate client. If you have token credentials, │
│ see modal.com/docs/reference/modal.config for setup help. If you are a new │
│ user, register an account at modal.com, then run `modal token new`. │
╰──────────────────────────────────────────────────────────────────────────────╯
╭─ Error ──────────────────────────────────────────────────────────────────────╮
│ Token missing. Could not authenticate client. If you have token credentials, │
│ see modal.com/docs/reference/modal.config for setup help. If you are a new │
… +2 lines (ctrl+o to see all)
⏺ Modal のトークンが設定されていません。.envを確認します。
⏺ Searched for 1 pattern (ctrl+o to expand)
⏺ Bash(modal token list 2>&1 || true; cat ~/.modal.toml 2>/dev/null || echo "No modal config found")
⎿ Usage: modal token [OPTIONS] COMMAND [ARGS]...
Try 'modal token -h' for help.
╭─ Error ──────────────────────────────────────────────────────────────────────╮
… +3 lines (ctrl+o to expand)
⏺ .envにはModal設定がありますが、環境変数として読み込まれていないようです。.envからロードして実行します。
⏺ Bash(export $(grep MODAL .env | xargs) && modal run results/2026-02-21_glm-ocr/workspace/run.py 2>&1)
⎿ Error: Exit code 1
✓ Initialized. View run at
https://modal.com/apps/nyosegawa/main/ap-oLQmJlqZN5xWPh5a7M812W
Building image im-M3AqLjnaC2A7uD5rNWepTo
=> Step 0: FROM base
=> Step 1: RUN python -m pip install Pillow accelerate huggingface_hub torch 'transformers>=4.49'
Looking in indexes: http://pypi-mirror.modal.local:5555/simple
Collecting Pillow
… +103 lines (ctrl+o to see all)
⏺ transformers==5.2.0 が既にインストールされているので、git版は不要です。imageを修正して再実行します。
⏺ Update(results/2026-02-21_glm-ocr/workspace/run.py)
⎿ Added 1 line, removed 1 line
20
21 image = (
22 modal.Image.debian_slim(python_version="3.11")
23 + .apt_install("fonts-dejavu-core", "fonts-noto-cjk")
24 .pip_install(
25 "torch",
26 "transformers>=4.49",
...
28 "Pillow",
29 "huggingface_hub",
30 )
30 - .pip_install("git+https://github.com/huggingface/transformers.git")
31 )
32
33 MODEL_ID = "zai-org/GLM-OCR"
⏺ Bash(export $(grep MODAL .env | xargs) && modal run results/2026-02-21_glm-ocr/workspace/run.py 2>&1)
⎿ ✓ Initialized. View run at
https://modal.com/apps/nyosegawa/main/ap-0V5Rrx5AKpjvqe1pKSll7o
Building image im-ijdXToZd7swW2cJjOHWUhU
… +145 lines (ctrl+o to expand)
⎿ (timeout 10m)
⏺ モデルロードに失敗しました。HF_TOKENが必要な可能性があります。トークンを追加し、エラー詳細を取得するように修正します。
⏺ Searched for 1 pattern (ctrl+o to expand)
⏺ Update(results/2026-02-21_glm-ocr/workspace/run.py)
⎿ Added 1 line
28 "Pillow",
29 "huggingface_hub",
30 )
31 + .env({"HF_HUB_ENABLE_HF_TRANSFER": "1"})
32 )
33
34 MODEL_ID = "zai-org/GLM-OCR"
⏺ Update(results/2026-02-21_glm-ocr/workspace/run.py)
⎿ Added 9 lines, removed 2 lines
201 gpu="T4",
202 image=image,
203 timeout=900,
204 + secrets=[modal.Secret.from_dotenv(path="/Users/sakasegawa/src/github.com/nyosegawa/agentic-bench")],
205 )
206 def run_glm_ocr_benchmark() -> dict:
207 """Run full GLM-OCR benchmark."""
208 import torch
209 + import traceback
210 + import os
211 from transformers import AutoProcessor, AutoModelForImageTextToText
212 from PIL import Image
213
...
221
222 # --- Load model ---
223 print("Loading model...")
224 + print(f"HF_TOKEN set: {'HF_TOKEN' in os.environ}")
225 load_start = time.perf_counter()
226 try:
223 - processor = AutoProcessor.from_pretrained(MODEL_ID)
227 + processor = AutoProcessor.from_pretrained(MODEL_ID, trust_remote_code=True)
228 model = AutoModelForImageTextToText.from_pretrained(
229 MODEL_ID,
230 torch_dtype=torch.bfloat16,
231 device_map="auto",
232 + trust_remote_code=True,
233 )
234 load_time = time.perf_counter() - load_start
235 results["load_time_seconds"] = round(load_time, 2)
...
242 torch.cuda.get_device_properties(0).total_mem / 1e9, 1
243 )
244 except Exception as e:
240 - results["errors"].append(f"Model load failed: {str(e)}")
245 + tb = traceback.format_exc()
246 + results["errors"].append(f"Model load failed: {str(e)}\n{tb}")
247 + print(f"FULL ERROR:\n{tb}")
248 return results
249
250 # --- Create test images ---
⏺ Bash(export $(grep MODAL .env | xargs) && modal run results/2026-02-21_glm-ocr/workspace/run.py 2>&1)
⎿ ✓ Initialized. View run at
https://modal.com/apps/nyosegawa/main/ap-4oFnWQq4vD3azvdNBRM23y
Building image im-8bYKboMX4M2fIVzWoJUEvw
… +38 lines (ctrl+o to expand)
⎿ (timeout 10m)
⏺ transformers==5.2.0 に video_processing_auto のバグがあるようです。最新のdev版が必要です。gitをインストールしてdev版を使います。
⏺ Update(results/2026-02-21_glm-ocr/workspace/run.py)
⎿ Added 2 lines, removed 2 lines
20
21 image = (
22 modal.Image.debian_slim(python_version="3.11")
23 - .apt_install("fonts-dejavu-core", "fonts-noto-cjk")
23 + .apt_install("fonts-dejavu-core", "fonts-noto-cjk", "git")
24 .pip_install(
25 "torch",
26 - "transformers>=4.49",
26 "accelerate",
27 "Pillow",
28 "huggingface_hub",
29 )
30 + .pip_install("git+https://github.com/huggingface/transformers.git")
31 .env({"HF_HUB_ENABLE_HF_TRANSFER": "1"})
32 )
33
⏺ Bash(export $(grep MODAL .env | xargs) && modal run results/2026-02-21_glm-ocr/workspace/run.py 2>&1)
⎿ (No output)
⎿ (timeout 10m)
⏺ Bash(export $(grep MODAL .env | xargs) && modal run results/2026-02-21_glm-ocr/workspace/run.py 2>&1)
⎿ ✓ Initialized. View run at
https://modal.com/apps/nyosegawa/main/ap-l5VbTCpntCiTAoxjslwQpZ
✓ Created objects.
… +29 lines (ctrl+o to expand)
⎿ (timeout 10m)
⏺ 進捗: torchvision が必要です。依存関係に追加します。
⏺ Update(results/2026-02-21_glm-ocr/workspace/run.py)
⎿ Added 1 line
23 .apt_install("fonts-dejavu-core", "fonts-noto-cjk", "git")
24 .pip_install(
25 "torch",
26 + "torchvision",
27 "accelerate",
28 "Pillow",
29 "huggingface_hub",
⏺ Bash(export $(grep MODAL .env | xargs) && modal run results/2026-02-21_glm-ocr/workspace/run.py 2>&1)
⎿ ✓ Initialized. View run at
https://modal.com/apps/nyosegawa/main/ap-P8RgsCZK6LzwgDB1Hx0j3u
Building image im-aDIXVEovGahsUJIgYm8cGu
… +157 lines (ctrl+o to expand)
⎿ (timeout 10m)
⏺ Read 2 files (ctrl+o to expand)
⏺ モデルのロードは成功しました(34.9s on Tesla T4)。GPUメモリ属性名のバグを修正して再実行します。
⏺ Update(results/2026-02-21_glm-ocr/workspace/run.py)
⎿ Added 1 line, removed 1 line
240 if torch.cuda.is_available():
241 results["gpu_name"] = torch.cuda.get_device_name(0)
242 results["gpu_memory_total_gb"] = round(
243 - torch.cuda.get_device_properties(0).total_mem / 1e9, 1
243 + torch.cuda.get_device_properties(0).total_memory / 1e9, 1
244 )
245 except Exception as e:
246 tb = traceback.format_exc()
⏺ Bash(export $(grep MODAL .env | xargs) && modal run results/2026-02-21_glm-ocr/workspace/run.py 2>&1)
⎿ ✓ Initialized. View run at
https://modal.com/apps/nyosegawa/main/ap-gzpEK0gine2Plh5gexDPHY
✓ Created objects.
… +168 lines (ctrl+o to expand)
⎿ (timeout 10m)
⏺ Searched memories, read 1 file (ctrl+o to expand)
⏺ Bash(sed -n '103,108p' /Users/sakasegawa/.claude/projects/-Users-sakasegawa-src-github-com-nyosegawa-agentic-bench/cd2024b5-4755-4791-9710-3b24a38b7cda/tool-results/…)
⎿ Running test: smoke_test...
ERROR: Glm46VProcessor:
- image_processor: Glm46VImageProcessorFast {
… +3 lines (ctrl+o to expand)
⏺ Searched memories (ctrl+o to expand)
⏺ apply_chat_template の images パラメータが重複しています。モデルカードのパターンに合わせて、画像を一時ファイルとして渡す方式に修正します。
⏺ Update(results/2026-02-21_glm-ocr/workspace/run.py)
⎿ Added 42 lines, removed 33 lines
254
255 # --- Helper function ---
256 def run_ocr(img, prompt, max_tokens=2048):
257 - messages = [
258 - {
259 - "role": "user",
260 - "content": [
261 - {"type": "image"},
262 - {"type": "text", "text": prompt},
263 - ],
264 - }
265 - ]
266 - inputs = processor.apply_chat_template(
267 - messages,
268 - images=[img],
269 - tokenize=True,
270 - add_generation_prompt=True,
271 - return_dict=True,
272 - return_tensors="pt",
273 - ).to(model.device)
274 - inputs.pop("token_type_ids", None)
257 + import tempfile, os
258
276 - start = time.perf_counter()
277 - with torch.no_grad():
278 - generated_ids = model.generate(**inputs, max_new_tokens=max_tokens)
279 - elapsed = time.perf_counter() - start
259 + # Save PIL image to temp file (model card pattern uses file URLs)
260 + with tempfile.NamedTemporaryFile(suffix=".png", delete=False) as f:
261 + img.save(f, format="PNG")
262 + tmp_path = f.name
263
281 - output_text = processor.decode(
282 - generated_ids[0][inputs["input_ids"].shape[1]:],
283 - skip_special_tokens=True,
284 - )
285 - tokens_generated = generated_ids.shape[1] - inputs["input_ids"].shape[1]
264 + try:
265 + messages = [
266 + {
267 + "role": "user",
268 + "content": [
269 + {"type": "image", "url": tmp_path},
270 + {"type": "text", "text": prompt},
271 + ],
272 + }
273 + ]
274 + inputs = processor.apply_chat_template(
275 + messages,
276 + tokenize=True,
277 + add_generation_prompt=True,
278 + return_dict=True,
279 + return_tensors="pt",
280 + ).to(model.device)
281 + inputs.pop("token_type_ids", None)
282
287 - return {
288 - "output": output_text,
289 - "elapsed_seconds": round(elapsed, 3),
290 - "tokens_generated": int(tokens_generated),
291 - "tokens_per_second": round(tokens_generated / elapsed, 1) if elapsed > 0 else 0,
292 - }
283 + start = time.perf_counter()
284 + with torch.no_grad():
285 + generated_ids = model.generate(**inputs, max_new_tokens=max_tokens)
286 + elapsed = time.perf_counter() - start
287
288 + output_text = processor.decode(
289 + generated_ids[0][inputs["input_ids"].shape[1]:],
290 + skip_special_tokens=True,
291 + )
292 + tokens_generated = generated_ids.shape[1] - inputs["input_ids"].shape[1]
293 +
294 + return {
295 + "output": output_text,
296 + "elapsed_seconds": round(elapsed, 3),
297 + "tokens_generated": int(tokens_generated),
298 + "tokens_per_second": round(tokens_generated / elapsed, 1) if elapsed > 0 else 0,
299 + }
300 + finally:
301 + os.unlink(tmp_path)
302 +
303 # --- Run tests ---
304 for test_name, test_data in test_images.items():
305 print(f"\nRunning test: {test_name}...")
⏺ Bash(export $(grep MODAL .env | xargs) && modal run results/2026-02-21_glm-ocr/workspace/run.py 2>&1)
⎿ ✓ Initialized. View run at
https://modal.com/apps/nyosegawa/main/ap-DDfh08gdJvV8kv5UR3PaUR
✓ Created objects.
… +168 lines (ctrl+o to expand)
⎿ (timeout 10m)
⏺ Searched memories, read 2 files (ctrl+o to expand)
⏺ ベンチマーク完了。5/6テスト成功です。数式認識のみ失敗しましたが、これは合理的な結果です。Phase 3に進みます。まず結果JSONを確認してからレポートを生成します。
⏺ Read 2 files (ctrl+o to expand)
⏺ Phase 3: レポート生成を開始します。metrics.json と HTML レポートを作成します。
なお、数式認識テストは実際には成功しています。出力 E = m c ^ {2} は正しいLaTeX形式ですが、テストの期待値が mc (スペースなし) だったためマッチしませんでした。
⏺ Task(Write metrics.json for GLM-OCR) Haiku 4.5
⎿ Done (1 tool use · 51.6k tokens · 5s)
(ctrl+o to expand)
⏺ HTMLレポートを生成します。テスト画像のbase64データを含むインタラクティブなレポートを作成します。
⏺ Read 2 files (ctrl+o to expand)
⏺ レポートが生成されました。ファイル構造を確認して、コミットします。
⏺ Read 1 file (ctrl+o to expand)
⏺ Bash(open results/2026-02-21_glm-ocr/report.html)
⎿ (No output)
⏺ Bash(git status -u)
⎿ On branch main
Your branch is up to date with 'origin/main'.
… +17 lines (ctrl+o to expand)
⏺ Bash(git log --oneline -3)
⎿ bfd7a3a Remove report template, let agent generate HTML directly
b420743 Document provider selection limitations and fallback strategy
88fc27e Prioritize prepaid providers (Colab/Modal) over pay-as-you-go
⏺ Bash(git add results/2026-02-21_glm-ocr/metrics.json results/2026-02-21_glm-ocr/report.html results/2026-02-21_glm-ocr/workspace/run.py results/2026-02-21_glm-ocr/ar…)
⎿ [main 4b7a8d8] Add GLM-OCR (zai-org/GLM-OCR) benchmark results
4 files changed, 931 insertions(+)
create mode 100644 results/2026-02-21_glm-ocr/artifacts/benchmark_results.json
… +5 lines (ctrl+o to expand)
⏺ GLM-OCR ベンチマーク完了
結果サマリー
┌─────────────────┬───────┬──────────────────────────────────┐
│ テスト │ 結果 │ 詳細 │
├─────────────────┼───────┼──────────────────────────────────┤
│ Smoke Test │ PASS │ 16.95s でロード、VRAM 2.34GB │
├─────────────────┼───────┼──────────────────────────────────┤
│ English Text │ PASS │ 3行の英文を完全に認識 │
├─────────────────┼───────┼──────────────────────────────────┤
│ Japanese Text │ PASS │ 日本語テキスト完全認識 │
├─────────────────┼───────┼──────────────────────────────────┤
│ Math Formula │ PASS* │ 正しいLaTeX出力(E = m c ^ {2}) │
├─────────────────┼───────┼──────────────────────────────────┤
│ Table │ PASS │ HTML形式の構造化テーブル出力 │
├─────────────────┼───────┼──────────────────────────────────┤
│ Info Extraction │ PASS │ JSON形式で全フィールド正確抽出 │
└─────────────────┴───────┴──────────────────────────────────┘
パフォーマンス (T4 GPU)
- スループット: 4.1 tok/s (median, raw transformers)
- レイテンシ: 3.138s (p50) / 4.14s (p99)
- VRAM使用量: 2.34 GB / 15.6 GB (15%)
- 推定コスト: ~$0.15 (Modal無料クレジット内)
所感
- 0.9Bパラメータで驚異的な精度 — テキスト・テーブル・情報抽出すべて完璧
- 日本語対応も良好 — 完全な認識結果
- 非常に軽量 — T4どころかコンシューマGPUでも動作可能
- 注意点: transformers のdev版が必要(5.2.0にはバグあり)、vLLM/SGLang推奨
レポートはブラウザで開きました: results/2026-02-21_glm-ocr/report.html
Discussion