Showing posts with label but. Show all posts
Showing posts with label but. Show all posts

Tuesday, February 21, 2017

Automating master failover is possible but needs care

Automating master failover is possible but needs care


I was asked from a few people about my opinion of the Githubs recent service outage. As a creator of MHA, I have lots of MySQL failover experiences.
Here are my points about failover design. Most of them duplicate with Roberts points.


- "Too Many Connections" is not a reason to start automated failover
- Do not repeat failover

I know some unsuccessful failover stories that "1. failover happens because master is unreachable (getting too many connections errors) due to heavy loads 2. failover happens again because the new master is unreachable due to heavy loads 3. failover happens again....". On database servers, newly promoted master is slower because of poor cache hit rate. On traditional active/standby environment, database cache on the new master is empty so youll suffer from 10x or even worse performance for the time being. On master/slave environment, slave has cache so performance is much better than standby server, but you cant expect better performance than master.
It does not make any sense to repeat failover within short time, and automated failover should not happen just because master is overloaded. If master is overloaded due to H/W problems (i.e. raid battery failure, disk block failure, etc), failover will need to be performed, but I think this can be manually done.

MHA does not start failover if specific error codes are returned (i.e. 1203: ER_TOO_MANY_USER_CONNECTIONS). And MHA does not repeat failover if 1. last failover failed with errors or 2. last failover happened within N minutes (480 minutes by default) ago.


- Do not failover if it is unclear master is dead

This is very important to avoid split brain. In many cases data inconsistency is more problematic than longer downtime. You need to make sure on the master that no mysqld process is running / will not run. Even though master is not reachable via TCP/IP connection attempts, mysqld may be just during crash recovery. Forcing shutdown on the mater (power off) is my favorite approach, but may take long time depending on H/W.
MHA has a helper script to kill (i.e. power off) master. When I developed MHA, I spent long time for investigating how to speed up shutting down machines.


- Prepare tools for manual failover

There are some cases that automating failover is really scary - typical example is a datacenter failure. If the whole datacenter is not reachable, it is not easy to automatically check masters status, and probably remotely shutting down master is not possible. And it would be unclear when the datacenter is recovered. In such cases I think automated failover should not be performed, but manual failover should be done. Proper alerts should be sent immediately, so that DBAs can start analyzing problems and start manual failover quickly. On master/slave environments, slaves relay log positions might be different each other. Checking all slaves status and if needed fixing consistency by parsing relay logs is painful. MHA will be helpful in such situations, and actually I have used MHA many more times for manual failover than automated failover.

Available link for download

Read more »

Monday, February 20, 2017

Back to basics in information security Proven year after year but apparently unattainable for many

Back to basics in information security Proven year after year but apparently unattainable for many


Im often wrong about many things in life...just ask my wife. However, Im feeling a bit vindicated regarding my long-standing approach to information security: address the basics, minimize your risks.

You see, more and more research is backing up what Ive been saying for over a decade. It what was uncovered in the new Cisco 2015 Annual Security Report. [i.e. "Less than 50 percent of respondents use standard tools such as patching and configuration to help prevent security breaches."...among many other things.]

The new Online Trust Alliance report found the same things.

CyActive had similar findings as well in their new study. ["Some of the worst attacks of this year could have been avoided, saving companies, governments and consumers millions of dollars" ... "Unfortunately, reactive defense remains the common denominator today, despite the overwhelming evidence of reused and recycled components seen in the most notorious attacks."]

The brand-new Trustwave State of Risk Report backs up this reality, and does so every year. [i.e. "60% run external vulnerability scans on critical systems (third-party hosted) less frequently than every quarter. Meanwhile, 18% never perform penetration tests." Ill venture to guess that 80+ percent of organizations are not looking at all of their systems that count...]

The same goes for the Verizon Data Breach Investigations Report.

Ditto for the Chronology of Data Breaches...on a daily basis.

These results combined with what I see in my work and Im even more convinced that if we focused on the basic principles of information security such as the ones I listed here six years ago, what I wrote for SearchSecurity.com in 2004, many of the concepts we learn during CISSP training...not to mention the ones listed in these two publications:
  • Security, Accuracy, and Privacy in Computer Systems, by James Martin (Copyright 1974)
  • Saltzer and Schroeder’s The Protection of Information in Computer Systems (Copyright 1975)

What gives!? Whats it going to take to fix our security problems?

No thanks, Ă˜bama, we dont need your approach to continued government growth thatll fix information security no more than your "healthcare" law has fixed healthcare.

Im not convinced we need a federal data breach law, either (thanks anyway, American Bankers Association). I believe we have enough laws on the books for now...

What were seeing in information security (i.e. people who ignore the basics and end up perplexed by why bad things keep happening) is not unlike what society does with social issues. Every generation has their own ideas on how to fix the worlds ills (namely passing more laws and redistributing more wealth) but were still not focusing the essentials that have proven to work across generations (i.e. free markets, lower taxes/regulation, coaching people to believe in themselves, etc.) and, thus, the problems continue.

As Jim Rohn once said: Success is easy, but so is neglect.

The title of this recent SC Magazine piece on the subject says we need a new approach. I respectfully disagree. We need discipline.

What we need to fix the security challenges we face are people willing to stop buying into the hype brought forth by vendors and analysts, especially those who stand to make money off of their shiny new products or services - not to mention the self-proclaimed security ninjas and cyber warriors who know everything about security, in their own minds. We then need these people to acknowledge that merely 20 percent of their security vulnerabilities are creating 80 percent of their problems. Finally, we need people who are willing to be leaders and step up do something about these weaknesses....

Otherwise, step aside and let someone else do what needs to be done.

I know its not that difficult. I see plenty of organizations who are successful in security. The problem is that most are not.

As Ayn Rand, author of Atlas Shrugged said, You can avoid reality, but you cannot avoid the consequences of avoiding reality. The time to start recognizing history and learning from other peoples woes is now. Use your power of choice...Dont be a dodger. Confront the issue, fix it, and get this behind you once and for all.

Available link for download

Read more »