I am using Quanteda in R and have created the corpus and dfm. However, I notice that the dfm and corpus contain less documents than the original file. I would appreciate if anyone could please let me know why this happens and how to fix? Thanks
0
votes
1 Answers
0
votes
You can try mention the docid_field and text_field explicitly something like this:
data_corpus = corpus(x = data,docid_field = "doc_id", text_field = "text")
where doc_id and text are columns in the dataframe data.
And then compute the Document Feature Matrix using dfm function of qunateda package
data_dfm = dfm(data_corpus)