I'm hoarding profile data diffs from a well known social network. Been crawling every single user for the past 2 years and saving the changes. Had to stop doing it last month, after storage costs became too much.
Can you share any of the techniques you use to avoid being blocked for scraping?
Interesting to hear that storage was too expensive. I'm sure you could find personal sponsors here who would help with funding in exchange for access to the data - even with a 'no commercial use' type of restriction!
I'd like to set it up as a subscription service like Domaintools. I'm also crawling and storing user's comments, friends, posts and links to media published in other groups.
It will make a great tool for profiling a person, but I'm not sure on how (il)legal it will be.