0
votes

I am using Quanteda in R and have created the corpus and dfm. However, I notice that the dfm and corpus contain less documents than the original file. I would appreciate if anyone could please let me know why this happens and how to fix? Thanks

1
Hey! welcome to StackOverflow! Please provide some context (e.g., code). It's difficult to help fix a problem if we don't know what is causing it :p - theforestecologist

1 Answers

0
votes

You can try mention the docid_field and text_field explicitly something like this:

data_corpus = corpus(x = data,docid_field = "doc_id", text_field = "text")

where doc_id and text are columns in the dataframe data.

And then compute the Document Feature Matrix using dfm function of qunateda package

data_dfm = dfm(data_corpus)