1
votes

We are trying to use Apache Cassandra for a large project and we have a Python script that runs INSERT queries into a Database Cluster.

While testing the script, on the developer laptops (MacOSX), it works perfectly and performs all INSERTs without a problem.

Every time it runs in production machines (Linux), it always has a:

cassandra.OperationTimedOut: errors={}, last_host=cassandra1.example.com

We are using the DataStax Python Driver and use more than one hosts (cassandra1.example.com and cassandra2.example.com) when creating the cluster.

Both computers have the same kind and level of access in terms of networks (no firewall, etc.). The production servers have 4 ms ping with the database while the developer laptops average at 40-50.

Any ideas what seems to be the problem?

2
It will help if you share your insert code. My gut says you're overwhelming the cloud servers which are probably less powerful than your osx box. - phact
Are you saying everything times out? Connections? Requests? All requests, or just some? - Adam Holmberg
Everything times out and nothing else is using the cluster. It's an empty Cassandra instance that I try to feed with data.. Same script on OS X laptop, in worse network conditions (10x RTT) and everything works normally.. - DaKnOb
Without more information, impossible to tell. What's the load look like on the servers when that happens? Is it always the same machine that fails? What's your Retry & Load balancing policy? What's your RF? How big of a cluster? Consistency level? Are you using batches? Are you doing async queries or sync? If async, are you using the cassandra.concurrent module or just blasting away? - Jon Haddad
I'm happy to answer any questions to solve the problem. The Cassandra cluster is a 3-node cluster that is brand new. That is, it has no data in it except from the table in which I'm trying to INSERT into. I have added all 3 machines in the Cluster initialization but always the first one is reported as last tried. The consistency level is default and untouched. I'm not using batches. I'm doing sync queries. Running the same script from a Mac OS X laptop works fine and results are added to the database. Running the queries by hand using cqlsh from the laptop and the cassandra servers - DaKnOb

2 Answers

0
votes

Most likely the 40-50ms network latency is slowing down the script enough that it doesn't overload the servers when run from your laptop. Where the production servers are closer so they can spam things much faster, and overload them. If you are spamming asynchronous writes as fast as possible, you probably need to throttle them back by checking the results every once in a while, or just doing rate limiting.

0
votes

The problem has been resolved and the solution is as follows:

We had a Python Class which was essentially used as an object. In the initializer, we created the cluster and connected to it and then passed the session/connection/cluster variables as properties using self.cluster, self.session, etc.

Later on, from a method in this Class, we called the execute() statement of self.session:

def executeQuery(self, id) self.session.execute("INSERT INTO table (id) VALUES (" + str(id) + ");")

We then replaced the object initializer with an empty function and put all Cassandra-related functions inside the executeQuery() method. The problem was resolved and no timeouts occurred.