Newspaper新聞社AI-OCR

Reading centuries of Japanese newspapers with AI-OCRAI-OCRで数百年分の日本の新聞を読み解く

1M+Pagesページ
1800sEarliest最古
1Knowledge graphナレッジグラフ
Archive → graphアーカイブ → グラフAI-OCR

Local Japanese newspapers hold a record reaching back to the late 19th century — much of it archaic and hand-written. The goal: make all of it searchable.日本の地方新聞には、19世紀後半まで遡る記録が残されています。その多くは古い手書き。目標は、それらすべてを検索可能にすることでした。

01The challenge課題

What they were up against直面していた課題

The archives reached back to early hand-written editions, with historical layouts, archaic characters, and inconsistent print quality across millions of pages. Standard OCR simply couldn't read them.アーカイブは手書きの初期版まで遡り、歴史的なレイアウト、古い文字、ばらつく印刷品質が数百万ページにわたります。標準的なOCRでは読み取れませんでした。

02How Solazu Holding helpedSolazu Holdingの支援

What we did私たちが行ったこと

Solazu Holding made a fragile, unreadable archive into a research tool anyone can search.Solazu Holdingは、脆く読めなかったアーカイブを、誰もが検索できる研究ツールに変えました。
  • Tuned a single AI-OCR model to recognize historical layouts and archaic characters across eras.ひとつのAI-OCRモデルを、時代をまたぐ歴史的なレイアウトと古い文字を認識できるよう調整。
  • Processed millions of pages, normalizing text so old and new editions live in one corpus.数百万ページを処理し、古い版と新しい版を同一コーパスにまとめるようテキストを正規化。
  • Structured the results into a knowledge graph linking events, people, and topics for instant search.結果を、出来事・人物・トピックをつなぐナレッジグラフへ構造化し、瞬時に検索可能に。
03The outcome成果

What changed変わったこと

Centuries of newspaper history became a searchable knowledge graph — letting researchers trace events and people across generations in seconds.数百年分の新聞史が、検索可能なナレッジグラフに。研究者は、世代をまたぐ出来事や人物を数秒でたどれるようになりました。

1M+Pages digitizedページをデジタル化
1800sArchives coveredからのアーカイブ
1Searchable graph検索可能なグラフ

Explore more case studies他の事例を見る

Want results like these?同じような成果を求めていますか?

Tell us what you're working on and we'll recommend the right solution — with a tailored quote.取り組み中の内容をお聞かせください。最適なソリューションと、貴社向けのお見積りをご提案します。