I have some pdf files that would make good longtail fodder. Anybody used a converter that wasn't crap? It doesn't have to be freeware.
Mechanical Turk?
That would be my last resort.
Well if nothing else works...
How many files? A lot I guess?
A metric crapload. Parts manuals, brochures, tons of them. The pdfs were just scans done a few years back, so the content is basically page after page of images.
I found one program that did break out each page into an individual jpg and then I could OCR it. That worked, but it's going to be too much work reassembling the pages in html.
http://pdftohtml.sourceforge.net/
I'm not so sure the text will end up as true html rather than simply images without OCR though
You could just breakdown and get yourself a Linux box RC.
>not so sure the text will end up as true html
Yeah, I'd seen that linux script and wondered the same thing. I'm thinking that it's talking about more recent types of pdf which were converted from Word docs or rtf.
>linux box
I'm sure I have one in a pile around here somewhere, hhh.
Fiverr.com?
>5r
looking more and more like it
i did try vsisoft today, it did spit out all the images numbered by the page it extracted them from. That could be handy. But the job is still a massive PITA. I'll probably trash the idea.
https://crocodoc.com/
Just read about it, then played with it - Looks very interesting
You also might be able to use: http://finereader.abbyy.com/
That came w/ my scanner and turns all the PDFs into searchable pdfs and does the OCR and what not. Then you could copy and paste.
Same with evernote. Not sure if the free one will do ocr or not, but you could then cut and paste.
This also seems like it may do a pretty nice job: http://www.verypdf.com/app/pdf-to-html-converter/index.html