I am using NLTK python to do sentiment analysis and my data has about 200,000 reviews. To use Naive Bayes Classifier, I need to have training set that is labeled. Since my data is not labeled, I manually created about 100 reviews as positive and negative. But I don't think this is the way to do it. I heard that I need to have 20% of data as a training set to train classifier and apply it to the rest 80% of data.
Is there any better way to generate training set for Naive Bayes classifier? Thank you for your help, and please let me know if the questions is not clear to understand.