# Bundled corpus source notices This prototype is not endorsed by any source publisher. Derived annotations and formatting changes are documented per source. No permission-specific UBS semantic fields are included. CC BY 4.0: https://creativecommons.org/licenses/by/4.0/ CC BY-SA 3.0: https://creativecommons.org/licenses/by-sa/3.0/ The morphological parsing and lemmatization from MorphGNT: SBLGNT Edition, by James Tauber and contributors, and Worden's adaptations of those annotations remain available under Creative Commons Attribution-ShareAlike 3.0 Unported (https://creativecommons.org/licenses/by-sa/3.0/). The SBL Greek New Testament text has its separate Creative Commons Attribution 4.0 International license (https://sblgnt.com/license/). A reusable annotation download matching this content release is available at https://wordenapp.com/sources. It includes the pinned MorphGNT source archive, Worden's occurrence/lemma/component identifiers, readable grammatical annotations, counts, attribution, checksums, and a standalone export tool. The download contains licensed data and the export tool; it does not include Worden's private application source. Other text and annotation sources retain their separately listed licenses. The independently licensed data may be copied, redistributed, and adapted under its applicable licenses. Worden does not apply its application-license restrictions to those rights. No source publisher endorses Worden. In the app, open Sources & licenses and use Export Greek annotations to save the matching unencrypted data and licenses without an internet connection. ## bsb The Holy Bible, Berean Standard Bible, BSB. Produced in cooperation with Bible Hub, Discovery Bible, unfoldingWord, Bible Aquifer, OpenBible.com, and the Berean Bible Translation Committee. Dedicated to the public domain. https://berean.bible/terms.htm Source: https://bereanbible.com/bsb.txt (third printing, snapshot 2026-09-06) ## hebrew-strong Open Scriptures HebrewLexicon / HebrewStrong.xml. James Strong, 1894. Public domain. Digitization and corrections by David Troidl and David Instone-Brewer. Source: https://github.com/openscriptures/HebrewLexicon ## macula-greek # MACULA Greek Linguistic Datasets [MACULA Greek Linguistic Datasets](http://github.com/Clear-Bible/macula-greek/) © 2022-2024 by [Biblica, Inc](http://biblica.com) is licensed under [CC BY 4.0 ](http://creativecommons.org/licenses/by/4.0/). These datasets include: 1. Greek Syntax Trees in both Node and Lowfat format 2. Morphology (`@lemma`, `@pos`, `@morph`, `@person`, `@number`, `@gender`, `@case`, `@voice`, `@tense`, `@mood`) 3. Word Senses (`../sources/Clear/wordsense`) distinct from the Louw-Nida Semantic Domains of Biblical Greek word senses 4. Semantic Frames (`../sources/Clear/annotations`) 5. Participant Referents (`../sources/Clear/annotations`) 6. Synonyms (`../sources/Clear/synonyms`) 7. Mappings (`../sources/Clear/mappings`) 8. Adjunct Types (`../sources/Clear/adjunct-types`) 9. Word Mapping between N1904 and SBLGNT: https://github.com/Clear-Bible/macula-greek/tree/sblgnt-trees/sources/Clear/mappings In addition to datasets from Clear, MACULA contains data from the following datasets. Where a repository is given, please refer to the repository for the relevant license: 1. [Nestle1904](https://github.com/biblicalhumanities/Nestle1904) Greek New Testament, edited by Eberhard Nestle, published in 1904 by the British and Foreign Bible Society. Transcription by Diego Santos, morphology by Ulrik Sandborg-Petersen, markup by Jonathan Robie. 2. [SBLGNT](https://github.com/LogosBible/SBLGNT) from Logos Bible Software; specifically, the data including the pericope adulturae [from their commit on July 10, 2023](https://github.com/LogosBible/SBLGNT/commit/736fdc76158950c3d04b949b7e013ca14305145a). 3. Word sense data from the United Bible Societies [MARBLE](https://semanticdictionary.org/) project. (`@ln`, `@domain`) Used with permission. 4. The [Berean Interlinear Bible](https://interlinearbible.com/) (`@gloss`). The Berean Bible and Majority Bible texts are officially placed into the public domain as of April 30, 2023. 5. **Cherith Glosses for the Greek New Testament**, by Andi Wu, Copyright (C) 2023 by Cherith Analytics, licensed under a Creative Commons Attribution 4.0 International License ("CC BY 4.0"). This information is represented in the `English` and `Chinese` attributes. # License ## Creative Commons Attribution 4.0 International (CC BY 4.0) This is a human-readable summary of (and not a substitute for) the [license](http://creativecommons.org/licenses/by/4.0/). ### You are free to: * **Share** — copy and redistribute the material in any medium or format * **Adapt** — remix, transform, and build upon the material for any purpose, even commercially. The licensor cannot revoke these freedoms as long as you follow the license terms. ### Under the following terms: * **Attribution** — You must attribute the work as follows: "MACULA Greek Linguistic Datasets, available at https://github.com/Clear-Bible/macula-greek/". You may do so in any reasonable manner, but not in any way that suggests the licensor endorses you or your use. **No additional restrictions** — You may not apply legal terms or technological measures that legally restrict others from doing anything the license permits. ### Notices: You do not have to comply with the license for elements of the material in the public domain or where your use is permitted by an applicable exception or limitation. No warranties are given. The license may not give you all of the permissions necessary for your intended use. For example, other rights such as publicity, privacy, or moral rights may limit how you use the material. ## macula-hebrew # Hebrew Linguistic Datasets [MACULA Hebrew Linguistic Datasets](http://github.com/Clear-Bible/macula-hebrew/) © 2022-2024 by [Biblica, Inc](http://biblica.com) is licensed under [CC BY 4.0 ](http://creativecommons.org/licenses/by/4.0/). These datasets include: 1. Syntax trees that combine the Westminster trees with OpenScriptures Hebrew Bible morphology, with a common set of identifiers. 2. Identifiers for orthographic words and morphs. 3. Greek equivalents drawn from the Septuagint 4. English and Mandarin glosses taken from Cherith's glosses. 5. Strong's numbers for both Hebrew and Greek equivalents 6. Semantic domains and word senses drawn from the Semantic Dictionary of Biblical Hebrew (see `@sdbh`, `@lexdomain`, `@coredomain`, `@contextualdomain`) 7. Semantic frames (see `@frame`). 8. Participant referents (see `@subjref`, `@participantref`) 9. Synonyms (see `./sources/Clear/synonyms`) 10. Word Senses (see `@SenseNumber`) that are distinct from the Semantic Dictionary of Biblical Hebrew word senses. In addition to datasets from Clear, MACULA contains data from the following datasets: - [Westminster Leningrad Codex](tanach.us/) - the somewhat informal [license](http://tanach.us/License.html) states that "All biblical Hebrew text, in any format, may be viewed or copied without restriction." The snapshot we used is in sources/tanach.us/xml. - [Westminster Hebrew Syntax without Morphology](https://github.com/Clear-Bible/macula-hebrew/tree/main/sources/GrovesCenter) Copyright (C) 1991-2018 by [The J. Alan Groves Center for Advanced Biblical Research](https://www.grovescenter.org/) is licensed under a [Creative Commons Attribution 4.0 International License ("CC BY 4.0")](https://creativecommons.org/licenses/by/4.0/). - [OpenScriptures Hebrew Bible](https://hb.openscriptures.org) is licensed under a [Creative Commons Attribution 4.0 International License ("CC BY 4.0")](https://creativecommons.org/licenses/by/4.0/). Original work of the Open Scriptures Hebrew Bible available at https://github.com/openscriptures/morphhb. See `./sources/OpenScriptures/xml` for our derivative version, which organizes morphology according to morph in order to make it easier to integrate into the trees. - [Semantic Dictionary of Biblical Hebrew](https://semanticdictionary.org/), edited by Reinier de Blois, with the assistance of Enio R. Mueller, ©2000-2021 United Bible Societies. Used with permission. - Cherith Glosses for the Hebrew Old Testament, by Andi Wu, Copyright (C) 2022 by Cherith Analytics, is licensed under a [Creative Commons Attribution 4.0 International License ("CC BY 4.0")](https://creativecommons.org/licenses/by/4.0/). See `./sources/Cherith/glosses`. # License ## Creative Commons Attribution 4.0 International (CC BY 4.0) This is a human-readable summary of (and not a substitute for) the [license](http://creativecommons.org/licenses/by/4.0/). ### You are free to: * **Share** — copy and redistribute the material in any medium or format * **Adapt** — remix, transform, and build upon the material for any purpose, even commercially. The licensor cannot revoke these freedoms as long as you follow the license terms. ### Under the following terms: * **Attribution** — You must attribute the work as follows: "MACULA Hebrew Linguistic Datasets, available at https://github.com/Clear-Bible/macula-hebrew/". You may do so in any reasonable manner, but not in any way that suggests the licensor endorses you or your use. **No additional restrictions** — You may not apply legal terms or technological measures that legally restrict others from doing anything the license permits. ### Notices: You do not have to comply with the license for elements of the material in the public domain or where your use is permitted by an applicable exception or limitation. No warranties are given. The license may not give you all of the permissions necessary for your intended use. For example, other rights such as publicity, privacy, or moral rights may limit how you use the material. ## morphgnt MorphGNT SBLGNT =============== [![DOI](https://zenodo.org/badge/1039950.svg)](https://zenodo.org/badge/latestdoi/1039950) Project to merge the MorphGNT analysis with the SBLGNT text. The SBLGNT text itself is subject to the [SBLGNT EULA](http://sblgnt.com/license/) and the morphological parsing and lemmatization is made available under a [CC-BY-SA License](http://creativecommons.org/licenses/by-sa/3.0/). How to cite ----------- Tauber, J. K., ed. (2017) _MorphGNT: SBLGNT Edition_. Version 6.12 [Data set]. https://github.com/morphgnt/sblgnt DOI: 10.5281/zenodo.376200 Columns ------- * book/chapter/verse * part of speech * parsing code * text (including punctuation) * word (with punctuation stripped) * normalized word * lemma Part of Speech Code ------------------- A- adjective C- conjunction D- adverb I- interjection N- noun P- preposition RA definite article RD demonstrative pronoun RI interrogative/indefinite pronoun RP personal pronoun RR relative pronoun V- verb X- particle Parsing Code ------------ * person (1=1st, 2=2nd, 3=3rd) * tense (P=present, I=imperfect, F=future, A=aorist, X=perfect, Y=pluperfect) * voice (A=active, M=middle, P=passive) * mood (I=indicative, D=imperative, S=subjunctive, O=optative, N=infinitive, P=participle) * case (N=nominative, G=genitive, D=dative, A=accusative) * number (S=singular, P=plural) * gender (M=masculine, F=feminine, N=neuter) * degree (C=comparative, S=superlative) NOTE: The part of speech and parsing codes were inherited from the CCAT tagging and will be deprecated in the next major release of MorphGNT. ## oshb # Open Scriptures Hebrew Bible This work is based on *The Westminster Leningrad Codex*, which is in the public domain. # License ## Creative Commons Attribution 4.0 International (CC BY 4.0) This is a human-readable summary of (and not a substitute for) the [license](http://creativecommons.org/licenses/by/4.0/). ### You are free to: * **Share** — copy and redistribute the material in any medium or format * **Adapt** — remix, transform, and build upon the material for any purpose, even commercially. The licensor cannot revoke these freedoms as long as you follow the license terms. ### Under the following terms: * **Attribution** — You must attribute the work as follows: "Original work of the Open Scriptures Hebrew Bible available at https://github.com/openscriptures/morphhb". You may do so in any reasonable manner, but not in any way that suggests the licensor endorses you or your use. **No additional restrictions** — You may not apply legal terms or technological measures that legally restrict others from doing anything the license permits. ### Notices: You do not have to comply with the license for elements of the material in the public domain or where your use is permitted by an applicable exception or limitation. No warranties are given. The license may not give you all of the permissions necessary for your intended use. For example, other rights such as publicity, privacy, or moral rights may limit how you use the material. ## tbesg TBESG - Translators Brief lexicon of Extended Strongs for Greek - STEPBible.org CC BY See also: TFLSJ - Translators Formatted full LSJ Bible lexicon 0-5624 - STEPBible.org CC BY TFLSJ - Translators Formatted full LSJ Bible lexicon extra - STEPBible.org CC BY ======================================================= Lexicons for the tagged texts used by STEPBible, based on BHS for OT and LSJ for NT, backwardly compatible for any tagging based on Strong numbers. Extended Strongs for Greek is backwardly compatible with original Strong and NASB, and extended to include all NT variants and LXX + variants. The Brief lexicon is based on the Abbott-Smith definitions, and is edited to conform with the extended Strongs. For a few words where Abbott-Smith lacks a definition, one is supplied from MiddleLiddel (MD) or STEPBible scholars The Full lexicon is edited from the Full LSJ by Tyndale House scholars, formatted to make it easy to read, with expansion of abbreviations and revealing examples from Greek literature on hover over dates of the earliest source (added by Tyndale House) ============================================================== Data created by www.STEPBible.org based on work at Tyndale House Cambridge (CC BY 4.0) ============================================================== This licence allows you to: * Include any part of this data in software or publications without requesting permission * Download the data and reformat it for your application, without changing the data * Send any proposed corrections to STEPBibleATGmail.com. to be verified (You MAY make changes yourself, but you should include a note of changes that can be viewed by those who use your new data) * Refer others to github.com/STEPBible as the source of the data. Please do not redistribute it yourself. (Updates or corrections are easier to implement when the data is distributed from a single source) * We'd love to hear about your project when you make it available. Email us at STEPBibleATGmail.com. ============================================================== ## Pronunciation conventions and sound key # Pronunciation candidates, recipe classroom-guides-1 The bundled corpus provides generated, form-specific sound guides. These are **unreviewed candidates**. The release quality gate stays closed until qualified Hebrew/Aramaic and Greek reviewers assess syllables, stress, and exceptional forms. Nonempty fields are not counted as reviewed pronunciation coverage. Greek uses an English classroom Erasmian convention, with accent treated as stress: ah (father), eh (met), ay (say), ee (machine), oh (go), oo (food), ue (French *tu*), eye (eye), oy (boy), ow (cow), kh (Bach), th (thin). Rough breathing is h; gamma before a velar is ng. Diphthongs count as one nucleus. Hyphens divide syllables, and capitals mark the accented syllable. Vowel quantity, pitch contours, historical changes, and regional alternatives are outside this recipe. Hebrew and biblical Aramaic use a simplified Sephardic classroom convention: ah, eh, ay, ee, oh, oo with kh (Bach), sh (ship), ts (cats); dagesh distinguishes b/v, k/kh and p/f. Shin/sin dots are retained, vowel letters and furtive patah are handled explicitly. Cantillation identifies candidate stress; headwords without accent marks are not assigned invented stress. Vocal/silent shewa, qamats qatan, guttural effects, matres, gemination and secondary accent require linguistic review. The vocalization written on the divine name is not silently replaced by an alternative reading. Its conventional liturgical reading is a pending exception requiring an explicit selected policy. The primary reading policy uses OSHB word elements outside notes (including ketiv as present in that source); qere notes are retained in the source archive, not silently substituted. It uses the complete MorphGNT primary text, which excludes John 7:53–8:11. Macula's differing edition is only an annotation source after an occurrence-level text check, never a replacement text. Reviewed corrections belong in `pipeline/pronunciation_overrides.json` with a source form, language, guide, reviewer and citation. No reviewed overrides have been asserted in this initial package. Current SBLGNT text license (checked 2026-09-06): Creative Commons Attribution 4.0 International, https://sblgnt.com/license/. This supersedes the old EULA wording still referenced in the pinned MorphGNT README. MorphGNT annotations retain their separate CC BY-SA 3.0 license. ## bsb-usx The Holy Bible, Berean Standard Bible (BSB). Dedicated to the public domain. Produced in cooperation with Bible Hub, Discovery Bible, unfoldingWord, Bible Aquifer, OpenBible.com, and the Berean Bible Translation Committee. Terms: https://berean.bible/terms.htm Source: https://bereanbible.com/bsb_usx.zip (publisher USX 3.1 snapshot, downloaded 2026-09-08). SHA-256: 7530e07b44b11ec42a2a6148d2707031dbafe8cdd616f2c383275f8106e79e4f Worden imports only explicit USX wj (words of Jesus) markup as UTF-16 ranges into the unchanged, separately pinned BSB text. Notes and headings are excluded. Layout whitespace is mapped to existing verse whitespace; wording and punctuation must match exactly. Nine reviewed chapter-ending literal [’’] artifacts are removed from annotation input only under a source-checksum-bound verse allowlist. Every reconciliation, unannotated non-speech mismatch, and absent USX verse is documented in red-letter-report.json. Other speech-bearing mismatches fail the import. This follows the publisher's editorial speech boundaries, including partial verses, without inferring speakers or projecting colors onto Greek or Hebrew. ## bsb-tables The Holy Bible, Berean Standard Bible (BSB). Dedicated to the public domain. Produced in cooperation with Bible Hub, Discovery Bible, unfoldingWord, Bible Aquifer, OpenBible.com, and the Berean Bible Translation Committee. Terms: https://berean.bible/terms.htm Source: https://bereanbible.com/bsb_tables.tsv (publisher translation tables, downloaded 2026-09-08). SHA-256: 09bbee6f9fe4fa22b5df28e8a9ffa99bf9c33435f4eb8c47c2dc221d855d35cb Worden imports the publisher's explicit English phrase cells, English order, native spellings and native order. Complete English verses must match unchanged pinned BSB text after documented table-formatting removal. Supplied words and untranslated placeholders remain unlinked. Snapshot-bound structural cleanup includes standalone vvv cells, optional opening/closing quote notation, and parenthesis-placeholder cells; every affected source row is audited. No Bible text is modified. Native occurrence matching uses only existing explicit passage correspondences to OSHB/SBLGNT, retains spelling, accents and vowel points, and permits only Unicode NFC, case and listed boundary-punctuation differences. A match must be shared by every optimal exact-token sequence alignment; ambiguous repeated words, textual variants, missing passages and missing correspondences remain unresolved. This does not infer word links from a lemma, dictionary gloss, verse number or semantic similarity. Phrase groups may cover multiple English ranges; supplied bracketed words remain unlinked rather than acquiring an inferred original-language word. Full coverage and unresolved source-row records are provided in word-alignment-report.json. ## Semantic encoder Model: sentence-transformers/all-MiniLM-L6-v2 Revision: 1110a243fdf4706b3f48f1d95db1a4f5529b4d41 License: Apache-2.0 Source: https://huggingface.co/sentence-transformers/all-MiniLM-L6-v2 Converted to Core ML with 16-bit weights; masked mean pooling and L2 normalization. Tokenizer vocabulary and matching BSB corpus vectors are packaged for offline use. The full model license is in semantic-model-LICENSE.txt.