AI RESEARCH
Best approach for OCR → cleaning → Hindi–English translation pipeline? [R]
r/MachineLearning
•
I’m working with a large CSV dataset (~200k rows) containing OCR-extracted text from parliamentary documents. The data includes a mix of Hindi (Devanagari) and English, often within the same sentence, along with OCR noise (broken words, symbols, formatting artifacts