Linked Clones Not A Panacea?

Thursday, December 11, 2008 Category : 7

Over at vinternals Stu asks if linked clones are the panacea that a lot of people are claiming about the storage problem with VDI? I say yes, however we are moving from designing for capacity to designing for performance, and VMware have given us some good tools to manage it. Let me explain a bit further.

Stu essentially raises two issues.

First, delta disks grow more than you think. Stu considers that growth is going to be a lot more than people expect, citing that NTFS typically writes to zero'd blocks before deleted ones and there is lots of activity on the system disk, even if you have done a reasonable job at locking it down.

Second SCSI reservations. People are paranoid about SCSI reservations and avoid snapshot longevity as much as possible. With a datastore just full of delta disks that continually grow, are we setting up ourselves for an "epic fail"?

These are good questions. I think what this highlights is that the with Composer the focus for storage for VDI has shifted from an issue of capacity management to performance management. Where before we were concerned with how to deliver a couple of TB of data now we are concerned with how to deliver a few hundred GB of data at a suitable rate.

In regards to the delta disk growth issue. Yes, these disks are going to grow, however this is why we have the automated desktop refresh to take the machine back to the clean delta disk. The refresh can be performed on demand, as a timed event or when the delta disk reaches a certain size. What this means it that the problem can be easily managed and designed for. We can plan for storage over commit and set the pools up to manage themselves.

To me the big storage problem we had was preparing for the worse case scenario. Every desktop would consume either 10G or 20G even though most only consumed much less than 10GB. Why? Just in case! Just in case one or two machines do lots of activity and because we had NO easy means of resizing them we also had to be conservative about the starting point. With Composer we can start with a 10GB image but only allocate used space. If we install new applications and decide we really do need the capacity to grow to 12GB, we can create a new master and perform a recomposition of the machines. Now we are no long building for worse case but managing for used space only. This is a significant shift.

So happens today there was a blog posting about Project Minty Fresh. This installation has a problem with maintaining the integrity of their desktops. As a result they are putting a policy in place to refresh the OS every 5 days. This will not only maintain their SOE integrity but also keep their storage overcommit it check.

In regards to SCSI reservations. I do believe that the delta disks do still grow at 16MB and not some larger size. So when the delta disks are growing there will be reservations, and you will have many on the one datastore. Is this a problem? I think not.

In the VMware world we have always been concerned about SCSI reservations because of server work loads. For server work loads we want to ensure fast and more importantly predictable performance. If we have lots of snapshots that SQL database system which usually runs fine now starts to behave a little differently. Predictability or consistency in performance is sometimes more important than the actual speed. My estimation is that desktop workloads are going to be quiet different. In our favor we have concurrency and users. All those users and going to have a lower concurrency of activity, given the right balance we should have a manageable amount of SCSI reservations, if not we rebalance our datastores, same space, just more LUNs. Also unlike servers, will users be able to perceive any SCSI reservation hits as they go about their activity. Given the nature of users work profile and that any large IOs should be redirected not into the OS disk but into their network shares or user drives the problem may not be as relevant as we may expect.

What Stu did not mention and we do need to be careful of because it can be the elephant in the room is IO storms. This is where we really do have some potential risk. If a particular activity causes a high currency of IO activity things could get very interesting.

Lastly, as Stu points out, statelessness is the goal for VDI deployments. Using application virtualisation, locking down the OS to a suitable level and redirecting file activity to appropriate user or networked storage is going to make a big impact on the IO profile. These are activities we want to undertake in any event, so the effort has multiple benefits.

I too believe you need to try this out in your environment, not just for the storage requirements, but also for the CPU, user experience, device capabilities and operational changes. VDI has come a long way with this release and I do strongly believe it will enable impactful storage savings.

What I really want is the offline feature to become supported rather than just being experimental. Plus I want it to support the Composer based pools. There is no reason why it can't and until then, there is still some way to go before we can address the breadth of use cases. However there are plenty of use cases now, which form the bulk, to sink our teeth into.

Rodos

Updated VMware Technical Resource documents listing

Category : 0

Two new documents added to the "VMware Technical Resource documents listing" on VMTN.

At version 16 there are now 197 documents listed with abstracts for searching.

Added Storage Design Options for VMware Virtual Desktop Infrastructure
Added Using IP Multicast with VMware ESX 3.5

http://communities.vmware.com/docs/DOC-2590

Rodos

VMware release searchable HCL system

Category : 0

VMware have released a new searchable HCL system that makes it much easier to check for compatibility. What does it look like and how do you use it? Read on.

Here is the URL
http://www.vmware.com/resources/compatibility/search.php
and the following image is an example search.



What do we have here? Lets say you have a BOM for a system, it includes a new card you are not familiar with, a NC360T. Enter that into the keyword search on the IOs tab. Great news, its supported in just about all versions, any they are all listed in front of you, could not be easier!

The page comes back with a number of components in the results. The first is a categorization based on partners. Lets say you type in something generic, this lets you quickly filter down to a particular vendor or subset, excellent feature.

For the results many of the details are hyperlinks. I have shown an exploded view of what it looks like when you click on a element. The opening page shows the specific details of that item.

This is a great new feature and is going to make our job so much easier. Do you self a favor and go and have a play with it, then create a bookmark!

Rodos

Is network acceleration useful for VDI?

Wednesday, December 10, 2008 Category : 2

Many of the WAN acceleration vendors (Cisco, Riverbed, Expand, Packeteer, Cisco WANscaler, Exinda) are sprouting some amazing statistics for improvement of VDI performance over a WAN. However are these claims realistic? How does one navigate the VDI and WAN acceleration space?

Here is a claim from one vendor

Expand’s VDI solution can can provide acceleration by an average of 300% with peaks of up to 1000% for virtual desktop user traffic.
Is this realistic? Sounds like a lot of sales and marketing to me. I remember reading a white paper from one vendor sprouting their lead in the acceleration of VDI, which consisted of the argument that by reducing the bandwidth of the other protocols over the link VDI would magically get faster. Whilst true, if that’s all there is it’s not really an improvement on RDP is it? I am sure this was Riverbed but I can’t find the paper and its not on their website anymore. Instead on the page about virtualisation they talk about ACE in the desktop space, good grief, get with the program, no mention of VDI at all.

So how do we make some sense of this space?

Firstly there are two areas that need addressing, RDP and printing.

Lets start with RDP, it’s a protocol that could do with some assistance over the WAN.

RDP as a protocol is chatty, by which I mean there are lots of smaller size packets which go back and forth, rather than a large data steam that goes generally in one direction like a file transfer. This makes RDP more affected by latency and packet loss. It needs some bandwidth but not huge amounts for normal workloads. As a guide I always use 30k to 100k bps per user with 150ms or lower latency. The bandwidth is going to depend on screen size, colour depth and the type of work you are doing. Also remember that as you add users its not a simple multiplier of bandwidth required, as you are more concerned about concurrency than number of users, not everyone is banging on their keyboards and refreshing their screens constantly or at the same time.

The WAN acceleration providers have three basic techniques for improving RDP itself.
  1. Improving the delivery and flow of TCP/IP. Riverbed call this Transport Streamlining and Cisco call it TCP Flow Optimization (TFO). Through lots of clever tricks at the TCP/IP layer to minimize effects of latency, fill packets, adjust window sizes and improve startup times these techniques can have a good impact on perceived and measured performance. The good thing here is that this occurs irrespective of the application protocol being transported, but different protocols will give different results.
  2. Compression. All the vendors use compression to reduce the amount of traffic. However the important point here is that if the traffic is already compressed or if it’s encrypted the effectiveness of this compression may be negligible. 
  3. Prioritization. If there are different protocols or even different sources or destinations, certain traffic can be prioritized over others. By jumping the queue time sensitive protocols such as RDP can show improved performance from the user perspective. However it’s important to remember that if the majority of traffic is all the same priority, improvements are not going to be achieved. If all the traffic over the WAN link is now just RDP traffic, it’s hard to make any improvements via prioritization.
There is another technique that the vendors sprout, which is application protocol specific acceleration, such as for CIFS or HTTP. By understanding the protocol and placing in optimizations to handle it specifically, dramatic gains can sometimes be achieved. However I have yet to see a vendor who when you dig deep, actually has techniques for improving RDP specific to its protocol. For RDP, it always falls back to the above categories of generic improvements around flow, compression and prioritization.

So how can this be put to advantage in VDI deployments? Here are my tips to give your WAN accelerator every chance at success.

  • If you do have a mix of other traffic over the same link you may get some good results. With other traffic bandwidth reduced and prioritization applied you should see some improvement in your RDP traffic.
  • Remove encryption from the RDP sessions. For your XP desktops you want to be installing the Microsoft patch http://support.microsoft.com/KB/956072 to allow the registry changes for turning the encryption level of RDP down or off. Of course you will need to take into account the security considerations and you may use some techniques in your WAN accelerator to add some encryption back in (after the fact). However removing the encryption allows for much better compression and caching of the data.
  • Turn off compression of the RDP sessions. This is described on page 152 of the View Manager 3 manual, the setting is “Enable compression” and it’s enabled by default.
Now let’s turn our minds to printing.

Moving to server based computing and therefore moving print jobs from the LAN to the WAN often causes headaches. Print jobs have a tendency to become large, which is not usually a problem when there is a 10 Mb or 100Mb link between the users machine and the printer. However put a number of users behind a much smaller link and print jobs start queuing up. The printer can sit their slowly receiving a large job over the small link, and nothing else prints until its all finished. That one small job of a single page right at the end of the queue has to wait and wait, as does the user who printed it. This is nothing new, it has been an issue in the Citrix world for a long time and vendors like ThinPrint have come out with some fantastic technologies to compress and prioritize print jobs, with some amazing results.

Where does WAN acceleration come into the printing issue? Today many WAN acceleration vendors include specific printing acceleration services within their solutions. Yes, VMware View Client for Windows comes with a version of ThinPrint but that is only for printing to a local printer through the client. This can provide part of the solution but most Enterprise customers in my experience will have networked printers in the remote offices run off central or remote servers, and with VDI the print jobs will not go via the client. Therefore you may find that your WAN acceleration solution also solves much of your new printing problems too. It may not provide a Universal Print Driver or have all of the nice features like prioritization (or it may) but it may be good enough.

So don’t be fooled by all of the marketing by the WAN acceleration vendors, they are all attempting to jump on the VDI bandwagon and it is causing some confusion. Try to understand where the improvements are actually happing, are they a primary improvement to the actual protocols or are they simply secondary effects of fixing other data going over the link? How will you be able to integration printing? Armed with some knowledge you should be able to evaluate vendors products based on their presented information plus also be able to conduct some good testing.

I plan to get some tests done with Cisco WAAS and VDI over the Christmas period if I can, I will let you know how I get on.

On a side note, the VMTN Community Round Table this week has Jerry Chen, Senior Director of Enterprise Desktop Virtualization as a guest this week. If Jerry runs out of things to talk about lets see if we can get him to give a view on acceleration of RDP and what he knows or has seen in the market? Maybe even some futures stuff or who VMware is working with.

Rodos

Does VMware have too many locations for technical materials

Wednesday, December 03, 2008 Category : 1

Do you think that VMware have too many locations of important or relevant technical materials? Its starting to feel there are a lot of places where some good content has the potential to be isolated or fragmented.

Here are some of the places that I know of just off the top of my head which contains content.

  1. VMTN. Here you will find community documents but some VMware teams put their content here too, such as the performance group.
  2. VMware.com Resources - Technical Resources page. This page lists all of the technical papers, probably the most valuable resource.
  3. VMworld Community. This is a new space and vendor and VMware information is starting to appear scattered throughout.
  4. VMware Whitepapers page. More general white papers, and this does link to the Technical Resources list.
  5. VI:OPS. the Virtual Infrastructure Operations site.
  6. Documentation sets.
  7. Knowledge base. Certainly a specific type of content thats going to be seperate. 
There are most likely others and I know there are some more on the way. I am sure there is an argument that each of these is targeted to a particular audience and its great to see VMware using the comminity to expand and enhance IP in multiple ways. However it does feel like things are starting to get fragmented and there is certainly a lot to keep your eye on if you want to be on top of the space. 

I wonder if VMware have a strategy here or are we seeing simply seeing Web 2.0 at work?

Rodos

Powered by Blogger.