Question about translator output vs “natural English” + lexicon versions #5
Unanswered
Alien-Tech-Solutions
asked this question in
Q&A
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Hi! Thanks for publishing this repo; it’s a really interesting project and the tooling is super helpful.
I was able to reproduce the translation pipeline pretty closely using the provided Python script plus the master/corpus JSON. I had a few clarification questions so I’m sure I’m using the most up-to-date sources correctly:
Are there any newer lexicon/corpus files than what’s currently in the repo (or a preferred branch/tag to use)?
Is the intended flow: script output = literal/systematic, and “natural English” = a condensed/collapsed rewrite of that output?
Are there any known lines/tokens that remain partially untranslated, or places where interpretation is still “in progress”?
Are there any personal methodology steps you use that aren’t encoded in the script (e.g., normalization rules, token-variant handling, special-case disambiguation)?
If it helps, I can share a short list of the small discrepancies I ran into (mostly around normalization / token variants), but I’d love to understand your “authoritative” workflow first.
Regardless if this is sitting repository for evidence or active, thank you for legitimately contributing and actually provide the ability for others to legitimately decipher the voynich manuscript with the same tools/knowledge and not providing false claims, a hoax, or nonsense while backing it all up with your own documentation.
All reactions