0
votes

I have a data frame with several million points in it - each having two values.

When I plot this like this:

 plot(myData)

All the points are plotted, but the plot is quite busy, so I thought I'd plot it as a line:

 plot(myData, type="l")

But while the x axis doesn't change (i.e. goes from 0 to 7e+07), the actual plotting stops at about 3e+07 and I don't actually get a proper line plot either.

Is there a limitation on line plotting?

Update If I use

 plot(myData, type="h")

I get correct and useable output, but I still wonder why the type="l" option fails so badly.

Further update

I am plotting a time series - here is one output using type="h":

Time series plot

That's perfectly usable, but having a line would allow me to compare several outputs.

2
You really need to plot millions of points? It will be even busier if you plot a line, because every point location will be connected by a line. If you really want to plot every point, how about just reducing the point size. You can do this with the cex parameter, e.g., plot(myData, cex=0.3). - eipi10
or, leave the point size to what it was, and set a highly transparent alpha value so that denser regions pop out, col="#00000022" usually works well. - Benjamin
The points follow a relatively smooth curve, so a line makes perfect sense - it's a time series - adrianmcmenamin
And this only happens when your data frame is big? Have you tried halving the size until it stops happening? Might give us a clue. Also, a reproducible example might help... - Spacedman
FWIW, I just plotted a ten-million point time series and didn't have a problem (except that it took a minute or so to render). Here's the code: set.seed(4943); dat = data.frame(time=1:1e7, values=cumsum(rnorm(1e7))); plot(dat$time, dat$values, type="l"). - eipi10

2 Answers

0
votes

High dimensional data graphic representation is growing issue in data analysis. The problem, actually, is not create the graph. The problem is make the graph capable of communicate information that we could transform in useful knowledge. Allow me to present an example to produce this point, by considering a data with a million observations, that is, not that big.

x <- rnorm(10^6, 0, 1)
y <- rnorm(10^6, 0, 1)

Let's plot it. R can yes easily manage such a problem. But can we? Probably not. Afterall, what kind of information can we deduce from an ink hard stain? Probably, no more than a tasseographyst trying to divinate the future in patterns of tea leaves, coffee grounds, or wine sediments.

plot(x, y)

enter image description here

A different approach is represented by the smoothScatter function. It creates a density plot of bivariate data. There, we create two examples.

First, with defaults.

smoothScatter(x, y)

enter image description here

Second, the bandwidth was specified to be a little larger than the default, and five points are specified to be shown using a different symbol pch = 3.

smoothScatter(x, y, bandwidth=c(5,1)/(1/3), nrpoints=5, pch=3)

enter image description here

As you can see, the problem is not solved. Nevertheless, we can have a better grasp on the distribution of our data. This kind of approach is still in development, and there are several matters that are discussed and evolved. If this approach represents a more suitable approach to represent your big dataset, I suggest you to visit this blog that discuss throughfully the issue.

0
votes

For what it's worth, all the evidence I have is that is computer - even though it was a lump of big iron - ran out of memory.