- Package:
- tesseract-ocr
- Source:
- tesseract
- Description:
- Tesseract command line OCR tool
- Submitter:
- Date:
- 2018-12-03 19:54:05 UTC
- Severity:
- normal
Legacy engine provides some data which are not available LSTM, e.g. font information, so it is not obsolete and a user should be able to use it regularily. The data for the legacy engine are in different directory, so there is no conflict. However the current package is configured to see a conflict here, e.g.
Hi Janusz, Tesseract 4 uses tesseract-ocr-deu and tesseract-ocr-script-frak. Tesseract 3 uses tesseract-ocr-deu-frak I am worried about confusing users. If we include both sets of language data in Debian, there will a huge number of choices, and some users might feel overwhelmed. However, Alexander thinks it could work with careful package naming, like this: https://mentors.debian.net/package/tesseract-lang-legacy I am also a little worried about the support costs of exposing lots of users to the legacy engine. It will make it harder to remove the legacy engine completely from future Tesseract. Especially if other Debian packages start to have dependencies on it. Also bug reports against the legacy engine will not get much attention from upstream. We'll discuss this with upstream, but in the meantime I have a question for you: What is your best guess for how many people are like you, and want to use the Tesseract 3 engine in Debian? Thanks, Jeff
I don't use Tesseract actively at the moment but subscribe to the tesseract issues. I don't read them carefully but have an impression that, at least at the very moment, Tesseract 4 traning data are not necessary better then the legacy ones. [...] To say the truth, I'm aware of only one other person. He became confused by the different paths to training data on Ubuntu and reported his problem as an issue, which has been closed almost immediately; unfortunately I'm unable to find quickly the issue numeber. Best regards Janusz