Friday, February 6, 2009

SQL Cluster configuration and moving the MSDTC.

So over the last 5 or 6 days we have been dealing with a mess of a SQL cluster here at the office. After speaking at length with MS on the issue we needed to perform the following actions.



Move the entire SQL Cluster group along with the MSDTC resource to the cluster group.

Fix the dependencies on the MSDTC resource.

Add a new clustered instance of SQL that can fail-over between all 3 nodes.



According to MS having the MSDTC resource depending on the SQL drive is NOT a supported configuration. MS SQL Team and the MS Windows Cluster Team both had to be involved in case things went bad during the move. MS wouldn't allow me to perform these cluster changes myself due to some possibility of the whole thing going south.



So we moved the MSDTC and SQL group. Then deleted the MSDTC and created a new one. Then we deleted the folder on the SQL data drive that had been there. We made the MSDTC dependent on the Quorum drive. So at this point MSDTC was all set. From there we moved all the SQL stuff to the right place in the SQL cluster first instance.

Then we installed the second instance and then upgraded it to fail over between all 3 nodes of the cluster.

Now that all the cluster stuff was fixed we had to move TEMPDB and the SQL DB and Tlogs and Backups. From there we were good. Since we were on an EMC array PowerPath FULL was then installed on all 3 nodes to ensure that connectivity via fibre-channel was as reliable as it could be.

Since it was 5am at this point and I had been working all night I calmed down and wrapped everything up. Then as quickly as it started it was all over and I could get some sleep :). LOL

Wednesday, February 4, 2009

Be BOLD!

In my employment history I have run into a number of situations that I wanted to discuss with other professionals. Mainly because true SPARKs or TALENTs understand and appreciate being in similiar types of situations.

I started asking myself the following questions:

  • If I could go back in time and speak to myself 10 years ago, what would I tell myself?
  • Would I discuss world-views and big picture corporate understanding?
  • What tidbits that I know now would I tell myself?
These are all good questions. If you are a young person or really ANY person that wants a deeper understanding on how to circumvent the "Old Garde" then keep your eyes peeled.

Go to : The Corporate Bold Homepage.

There you will read about a book that could help many people. I have been selected as a co-author. The book is slated to come out this summer. It is a collection of 100 short experiences. I can't give you anymore information, but realize that it's 100 experiences and pieces of truth from 100 Top Performers.

If you have questions visit the site and ask them.

More SQL Cluster stuff to come soon. :)

Monday, February 2, 2009

SQL Cluster Storage Changes.

This past weekend we had a client environment that needed to move from Microsoft iSCSI to Fibre for SAN connectivity. The reasons for this are numerous.

  • Fibre connectivity has greater throughput.
  • The Microsoft iSCSI initiator is junk for a clustered environment. ( when installing 2.07 the cluster nodes would intermittently lock up and lose resources, this is a known issue with 2.07. We updated to 2.08 and still had issues.)
  • PowerPath on Fiber has proven to be solid numerous times.
  • Less complexity.
  • Greater reliability.

We ran into some issues though and since I didn't see any direct posts on the web or on TechNet. So I will detail the issue and the solution here.

Specs:

Servers HP DL380G5; Windows 2003 R2 Enterprise 32-bit, SQL Enterprise 32-bit. Current cluster configuration is Active/Passive. Initial storage configuration is iSCSI. Moving to Fibre HBAs.

Issue:

After installing a Fibre HBA Dual port card and zoning it to the storage. We uninstalled the MS iSCSI initiator and disabled the iSCSI NIC ports on the server. Upon rebooting we attempted a full failover. The cluster group would fail over with the Quorum but the SQL instance would fail over the IP resources and the T log drive then hang. The result is a full failover is not possible, which in an environment that HAS to be online is not a good thing. The event ID would reference not being able to flush the transaction log.

Solution:

After reboots and trying to see if we missed anything these are the steps that we followed to resolve the issue:

  • Uninstall from windows the NICs used for iSCSI.
  • Remove the physical adaptor used for iSCSI from the server.
  • Reboot.

Once these steps were completed a full failover was again possible.

If you have any questions/concerns or if you are having this issue yourself please feel free to comment. We will do our best to assist you.

CC

Sunday, January 11, 2009

1TB 2.5" SSDs and 4TB 3.5" SSDs? SWEET Here we go!

http://www.marketwire.com/press-release/Puresilicon-936099.html

In the link above Pure Silicon has announced something TRULY usable in an SSD. Size and Speed. Size and Speed larger then the current drives. As prices come down the death of the rotating magnetic plate is inbound.

2 of these in a blade server running ESX has enough storage and performance to go up against small SAN installs.

-OUT

Wednesday, December 17, 2008

VMware without a Console. COS-Less and it's effect on you.

So in talking with some of the other professionals in the field we have discussed what needs to happen to make COS-less ESX the next big thing. ESXi is the first attempt and it's a good start. But there is a way to get the console in ESXi. In order to move this forward, all the console functionality will need to put in vCenter or something.

I was just thinking how Sweet it would be if when they move completely COS-less if you came up to the ESXi splash and entered in all the information and then put in the address of the vCenter server if you had one.

Once COS-less ESXi connects to the vCenter server you could have the options to roll-out a specified config (ala answer files for custom roll-outs) or push down a base config to use later. The only information you would need is the IP information on the hosts.

I can see a number of situations where this type of controlled roll-out would be beneficial. The other issue running through your head right now is "What if I don't have a vCenter server?". It's a good one. Maybe a limited functionality config database could reside on a server called "basic vCenter" some type of free version. By the time that this would happen clustered vCenter servers would have become the norm so up time would be slightly less of a concern.

I don't know if this is the way that VMware is heading towards but it would be cool nonetheless.

Let me know your thoughts. Then we can run through pro's and con's on the next blog article.

Friday, December 5, 2008

Building an Enterprise Class VMware Infrastructure. Take your time, blowing it here could cost you.

Designing an Enterprise class VMware VI3 environment is not an incredibly difficult task. It is though one that takes good planning and a full understanding of the network in both your data center and how your internal processes work. You also need a fair amount of VI3 understanding. You'll need to do your homework. A bunch of homework. So be ready. Remember that Enterprise means production, so treat it that way.

In the next few paragraphs I'll go through some of the items that I find it critical to look at when beginning a design.

So let's go through some of the basics.


1) Back-End Storage.


-You might be thinking "Why is storage given such a prominent place in design?". Here's your answer - Everything rides on your storage. Therefore it has to be more than adequate. You have to know your Read/Write ratio the number of IOPS you need to have available for the hosts, rough growth estimates, and don't forget adequate space.

So how do you know what to do?

Take valid measurements from physical hosts and plan your design around them. Example : Most VMFS volumes have a Read/Write ratio of 75/25. That makes them perfect for RAID 5. If your Read/Write moves more towards 50/50 or higher RAID 10 becomes a need instead of a desire.

Let's not discuss the "pooled storage" concept and what needs to be done to rid the market of it's presence. This is going to be long-term production and critical, treat it that way. Dedicated RAID sets for VMFS volumes. Let's remember the configuration maxims for VMware 64VM's per LUN is the MAX but you want to keep it to 20 or so as a sweet spot. Size your VMFS properly, if you plan to use a total of 800 GB remember that VMware needs some space to play just like Windows and Linux. That means your VMFS should be about 1TB in size.

Now about performance, you do know how much IO your going to push at these volumes if you did your homework. Please do yourself a favor and use at least SAS or FC-AL disks instead of SATA. We spoke earlier about VMware's intended purpose. It is production so please don't cheap it out with SATA. If you want to absorb 1200 IOPS you can't just use the RAW numbers of Disk performance. You have to calculate out your needs. Let's use a 75/25 Read/Write ratio and see what kind and how many disks we need to meet 1200 IOPS of performance on the front-end. To calculate the back-end storage we can use this formula for RAID 5 = (Disk IOPS * Read Ratio)+((Disk IOPS*Write Ratio)*4) ... this would give us a total of 2100 RAID adjusted IOPS. Therefore we will need at least 13 Disks in a RAID 5 array that spin at 15K (assuming we get 170 IOPS out of a 15K disk). So in the EMC world the best option would probably be 3 raid sets using a 4+1 Raid 5 layout and doing a metaLun across all of those. You would get some space and a touch of extra performance due to the metaLun. So do all that work for every VMFS or LUN that you need, don't forget to account for growth. In doing this you will properly size your storage environment and not have to go right back to the "well" in order to get more storage because performance sucks.

2) Network Infrastructure.


Networking in ESX is going to be more robust with the Nexus1000V but since we don't have that option yet, let's plan on the real networking horsepower to be in your core. I.E. - Cisco 6500, Foundry "JunkIron" ;P or your other various flavors of the "Core". If you need to span a bunch of VLAN's make sure that you have the vSwitches set for your needs. My personal preference is Tagging the frames at the vSwitch level. Then sending the frames over to the "Core" on Trunked links. Make sure you account for the amount of network usage you need. Don't under size this as NIC's are not overly expensive. Don't forget vMotion and redundant Service Console NIC's as they will play into your total NIC count.

Also make sure that you and your network guy have gone over this closely. If you are the "everything guy" double and triple check yourself. Make sure that you can account for all the bandwidth you need plus growth and the inevitable traffic spikes.

Also make sure that you connect ESX into the "Core" properly. If you are planning on EtherChannel then make sure that you have IP hash set on the NIC teaming for the vSwitch and portchannels properly configured on the switch.

If you have an internal vSwitch to vSwitch implementation. Please remember to account to everything on the inside of it so that you are not overwhelmed.

3) Server OS.

Next thing that I like to check, the Server Guest OS mix that I am going to run. If this is a production environment and you are planning on doing P2V's for most of your VM's then this is not a big deal. I like to verify that .ISO's for all the OS flavors I need are in a dedicated VMFS store that has been provisioned and has good performance. This keeps anyone from spending time locating OS media which is a waste of time that you can avoid.


4) Policies.

Policies are more of a "Who can screw up what." discussion with the Admin team. Locking people who don't understand VMware OUT of the system might be a good idea. After all, how many times have you seen the "IT Manager" think he understands all the technology log in and junk a VM because he didn't know what was going on? (I have. It happened more than once.) Don't lock management completely out. Just don't give them the creds to "help" you. Always make sure that no ONE person has UBER power. Always have a check and balance.

5) VI Host Hardware.

Host hardware is one of the places where you can demonstrate strategic ROI-based thinking. Looking to have the company spend just enough and then when the time comes to expand all the quantities are known and you don't need to perform the ENTIRE design process all over again. Buying 4 huge 16 proc boxes might not be the best use of funds. But if you got a number of Dual and a few Quad proc boxes it gives you a certain flexibility that cannot be underestimated. Being strategic here will demonstrate to your boss and those around you that your worry is the whole road map and not just a single point on a single solution. You gain credibility in a number of ways.

Some Bullet Points for Host Hardware:
  • Choose Either All AMD or ALL Intel don't mix/match.
  • Choose ONLY hardware on the HCL.
  • Forward thinking pays off here so do some.
  • Blades or Pizza Boxes. Mixing the two is just plain dumb. *cough* CoGR *cough* (In the end you get saddled with not being able to benefit 100% from either technology).
  • Build around total "Pools of resources" instead of being worried about individual specs.
  • Use EVC so that you can have forward mobility in your deployment.


4) Goals for implementing VMware.

Make sure that these Goals are documented.
Don't document the Goal without documenting the metrics that apply to them.
Verify that you are on track with the deployment.
Provide actual numbers for management to see and contrast the differences.
Make sure that you have planned to exceed your targets. (I know it seems elementary but it helps to have it in the back of your mind.)

**Do you have feedback or would you like an area of this article expanded? Please let me know.**

Wednesday, December 3, 2008

VI Too Easy? Too Easy To Screw-UP! YES.

So people have begun to see the value proposition in VMware and virtualization in general. But goodness. It is so easy to screw up the config and have it still perform! That speaks to the strength of the platform though. So how do we as VCP's and IT professionals change the CIO's and CTO's from ordering the "underlings" to deploy this great new technology when they don't even have a clue? We provide "Thought Leadership". We can give the concepts. Phil over at Joe the Consultant brings up a number of good points in one of his early blog posts titled "Why, Whatever shall we do?". In these tough economic times it's tough to see the value in bringing in a consultant to "just install ESX". But that is the problem. As consultants CTO's and CIO's need to recognize that we bring a whole HOST of services. Understanding both the technology and it's management so that the "corner stone" of your computing infrastructure can be managed and maintained in a proper manner. Is it too costly to spend 3-5000$ in the early stages to build it properly or 4 times that once it's running and the internal staff has hit a brick wall?

Front-loading time into design and understanding ALWAYS pays off during implementation.

So if you as a CIO or CTO are thinking about moving "Full-Speed" into virtualization, then please consider bringing in a consultant and tapping into their experience before starting out. They can guide you in your needs for Network Connectivity, Storage, and overall build practices. That money you spend up front will be easily recovered in NOT making common mistakes. This consultant can then work with you to determine a training path for your internal staff in order maintain the vision. Build a road mapped solution. Point solutions will only last for a short time and in the end will not provide either your internal clients OR your external clients the needed service that you can deliver.