AI SeedbankHelp preserve open and free AI for humanity's future

← All models

tohoku-nlp_bert-base-japanese-whole-word-masking

tohoku-nlp · View on Hugging Face ↗

Get this model

Download TorrentMagnet Link

Seeders: · Leechers:

Model card

The complete upstream card, rendered from this payload's README.md — the same hash-verified bytes the torrent distributes. Images and off-site links are removed; the original card on Hugging Face carries them.


language: ja license: cc-by-sa-4.0 datasets:

  • wikipedia widget:
  • text: 東北大学で[MASK]の研究をしています。

BERT base Japanese (IPA dictionary, whole word masking enabled)

This is a BERT model pretrained on texts in the Japanese language.

This version of the model processes input texts with word-level tokenization based on the IPA dictionary, followed by the WordPiece subword tokenization. Additionally, the model is trained with the whole word masking enabled for the masked language modeling (MLM) objective.

The codes for the pretraining are available at cl-tohoku/bert-japanese.

Model architecture

The model architecture is the same as the original BERT base model; 12 layers, 768 dimensions of hidden states, and 12 attention heads.

Training Data

The model is trained on Japanese Wikipedia as of September 1, 2019. To generate the training corpus, WikiExtractor is used to extract plain texts from a dump file of Wikipedia articles. The text files used for the training are 2.6GB in size, consisting of approximately 17M sentences.

Tokenization

The texts are first tokenized by MeCab morphological parser with the IPA dictionary and then split into subwords by the WordPiece algorithm. The vocabulary size is 32000.

Training

The model is trained with the same configuration as the original BERT; 512 tokens per instance, 256 instances per batch, and 1M training steps.

For the training of the MLM (masked language modeling) objective, we introduced the Whole Word Masking in which all of the subword tokens corresponding to a single word (tokenized by MeCab) are masked at once.

Licenses

The pretrained models are distributed under the terms of the Creative Commons Attribution-ShareAlike 3.0.

Acknowledgments

For training models, we used Cloud TPUs provided by TensorFlow Research Cloud program.

Magnet link

Opens the swarm directly in your torrent client — no file download needed. Copy-paste works too:

magnet:?xt=urn:btih:e8bd8aea50088637c2f69766695089a6f9d19914&dn=tohoku-nlp_bert-base-japanese-whole-word-masking

Open magnet in torrent client · infohash e8bd8aea50088637c2f69766695089a6f9d19914

Files & hashes

PathSizesha1sha256
README.md2.1 KB (2,137 B)7d29b4880b0b0473acb27eb18e2f4c65f2a8e071a08086b01dada685b56a0640a73b5de158175e5ca1f7852cab5093faabe10de5
config.json479 B (479 B)8af97449d6fff1379988a1afa8ba92eff8c44744bee86ed8692ce85de6cae9fbecbd4f8da22bdf1e6c87553b803e3f489dfb77f6
pytorch_model.bin424.4 MB (445,021,143 B)ff70d156b795d3f407fde5b0b6d9efa72137ae086df02178db2694615edb50e85581f40c002a10ac6ca58d966eef921a6ac5a27a
tokenizer_config.json120 B (120 B)e559d2e6aa3ba7c4d15bdc663d76412b4ad331d548f1b70f33a7b0457850fbbb56ff40492234f9e20e3328d1a62ae517f1b2051f
vocab.txt251.7 KB (257,706 B)d82671b24323c0659d1fd163cb50979c341497ce4391030cfeb780a85e807f33da97044123eeb38517fb3b2bc811b82fe723384a

Cite this release

Canonical URL
https://aiseedbank.org/models/tohoku-nlp_bert-base-japanese-whole-word-masking/
Slug
tohoku-nlp_bert-base-japanese-whole-word-masking
Infohash
e8bd8aea50088637c2f69766695089a6f9d19914
License
cc-by-sa-4.0
Signing key fingerprint
85a3b32c3712427b

Every file carries a locally computed sha256 — verify a download against the signed sums: tohoku-nlp_bert-base-japanese-whole-word-masking.SHA256SUMS (+ minisign signature).

Provenance

Upstream repositorytohoku-nlp/bert-base-japanese-whole-word-masking
Revision (pinned)861c1efefb044607ff5d1af7b5bb4c500279c8ab
Fetched at2026-09-04T06:35:20Z
License at fetchcc-by-sa-4.0
Snapshot toolhuggingface · seedbank 0.1.0

Trackers

✓ verified · rehash-vs-hf-metadata at 2026-09-04T06:35:28Z

cc-by-sa-4.0424.7 MB (445,281,585 bytes)transformerspytorchjaxbertfill-maskendpoints_compatible2 languages (tf, ja)