Conversation
open source OCR is terrible, I used Google Cloud to translate and that worked a lot better.
3
0
0
@sun tesseract isn't great. there are qwen models and such that do ok.

GLM has one that can read handwriting BUT it requires you to have already pre-processed the input with another model :/
1
0
2
@sun aren't there ai models that can do it? Iirc reading text was the first thing anyone did with artificial neural networks
1
0
0
@snacks I tried about three different things and using the google cloud document api is the only thing that worked acceptably.
1
0
1
@sun @snacks you could try the gemma models, the small ones are really fast
1
0
0
@lain @snacks it cost like 150 yen to convert this 80 page document, totally worth it.
0
0
2
@sun
>open source OCR is terrible
>I used Google Cloud to translate
>to translate
What ? OCR isn't about translating.
All the free/libre OCR I know just work.
1
0
0
@mangeurdenuage translate to text, sorry that's confusing terminology but technically correct.
1
0
0
@icedquinn @sun
>tesseract isn't great
I must live in another dimension but for me it works flawlessly. the last time I had to use it and that was 3 years ago.
2
0
0
@sun
Have you tried the FF translation model ?
0
0
0