My Notes

Picture the Words

Sources & Licenses

Last updated: .

Picture the Words is built mostly from original work — the flashcard structure, the illustration direction, the example curation, and the site itself. A few pieces draw on third-party open data, fonts, and software. This page lists everything we've identified, with the license that actually applies and where we sourced it, so it's separate from and doesn't get buried inside the Terms of Use.

A note on precision: we verified each license below against its primary source rather than assuming. Where a resource's licensing history is genuinely unclear even at the source, we've said so rather than guessing.

1. Picture the Words original content

The flashcard set (word choices, translations, example selection, card layout and template, category structure), the site design, and the written content on this site are original to Picture the Words, © Picture the Words, except where a third-party source is credited below.

Flashcard illustrations are produced with an AI image-generation tool (OpenAI's image API), directed and curated by Picture the Words for each concept. They are not hand-drawn, and they are not stock photography or third-party artwork.

2. Dictionary data — JMdict/EDICT

The Picture the Words browser extension's word-lookup feature uses dictionary data from JMdict/EDICT, a Japanese–English dictionary file maintained by the Electronic Dictionary Research and Development Group (EDRDG).

3. Example sentences — Tatoeba Project

Japanese example sentences shown on flashcards come from the Tatoeba Project, a free collaborative sentence database.

4. Example sentences — Wiktionary (removed 2026-09-14)

Until 14 September 2026 a small number of flashcards carried an example sentence from the English Wiktionary, used only for words the Tatoeba corpus (§3) has none for, sourced via kaikki.org's structured extraction rather than scraped directly. Those 6,036 sentences have been removed.

5. Hindi and French dictionaries — FreeDict and WOLF

Looking up a Hindi or French word that isn't one of Picture the Words' own 5,000 concepts falls through to two open dictionaries, the way a Japanese word falls through to JMdict (§2).

How WOLF was built, and why it matters. Most of the sources on this page were compiled by people: JMdict by the EDRDG's editors, FreeDict's Hindi from Shabdanjali at IIIT Hyderabad, Tatoeba's sentences by its contributors, FLORES-101's by professional translators. WOLF was not. Its own paper is titled Building a free French wordnet from multilingual resources — it was produced by automatically mapping Princeton WordNet onto French using bilingual dictionaries and parallel corpora, with partial rather than exhaustive manual validation.

In practice that means a French word's English senses are usually right and occasionally not, and no individual entry carries a human signature. We checked the alternative: WoNeF is a later and better-evaluated automatic translation of the same WordNet, but it is automatic as well — its published gold standard covers roughly 900 synsets checked by two annotators, out of more than 115,000 French entries. No freely licensed French wordnet is hand-built throughout.

So the French fallback dictionary is automatically derived and is described that way here rather than presented as equivalent to the hand-compiled sources above. It is used only for words outside Picture the Words' own 5,000 concepts — those 5,000 were translated and reviewed by people, and nothing on this page changes that.

Not legal advice: the GPL was written for programs, and there is no settled answer on whether shipping a GPL-licensed dictionary file alongside a closed application makes the two a single combined work or merely an aggregation. We treat it as aggregation and attribute accordingly. If this matters for a business decision, get it checked by a lawyer before relying on it.

A note on the phrase feature. The phrases shown on a card — faire du vélo, साइकिल चलाना, 自転車に乗る for “to cycle” — are not drawn from any dictionary on this page. They are Picture the Words' own concept labels, translated and reviewed by native speakers along with the rest of the concept set, simply collected together so you can see where one language needs several words for what another says in one. No new source, and nothing automatically generated.

6. Tanaka Corpus

Picture the Words uses Tanaka Corpus data to help match example sentences to the word being studied (word indexing), rather than as the direct source of sentence text, which comes via Tatoeba (§3).

Flagged for review: this licensing history should be confirmed (or a decision made on how conservatively to treat it) — see the note at the end of this page.

7. Fonts

All typefaces are loaded from Google Fonts under the SIL Open Font License 1.1, which permits commercial use and does not require attribution in the product itself.

8. Audio

Spoken audio for flashcards is synthesized using Microsoft Azure AI Speech (male English and Japanese neural voices). This is a paid commercial cloud service, not open community data, and no public attribution is required for using it in a downstream product. No community-recorded or crowdsourced audio (e.g. from Tatoeba) is used anywhere in the product.

9. Software libraries

Payment processing is handled via Paddle and Razorpay's own checkout scripts, governed by their respective merchant terms of service rather than a content license, so they aren't listed as attributed content above.

A note on legal certainty

This page reflects attribution and licensing requirements we identified from the available source licenses at the time of writing, verified against each resource's own primary documentation where possible. It is not a substitute for legal advice, and one item above (the Tanaka Corpus) has a licensing history we could not fully resolve from primary sources. If you have information that clarifies it, or a concern about anything on this page, please get in touch.