Showing posts with label IaaS. Show all posts
Showing posts with label IaaS. Show all posts

Sunday, January 16, 2011

Google excludes scheduled maintenance from its Google Apps SLA, but not Google App Engine

Back in July of 2010, I wrote about how Cloud Service Providers exclude scheduled downtime from their service level agreements.  Last Friday, Google made a significant change to their SLA for Google Apps by removing scheduled downtime from their Google Apps SLA:
  • Exclusion of scheduled downtime from availability SLA
  • Exclusion of intermittent downtime (periods of less than 10 minutes) from availably SLA
Obviously, it is good news for Google Apps customers.  It highlights Google’s infrastructure & operational maturity (in this case, for Google Apps specifically).  

I also think the announcement is important, because it sets a higher standard for service delivery.  By raising the bar, Google also intensifies competitive pressure on service providers such as Microsoft to offer more robust Cloud services.  Ultimately, both customers & the industry should benefit from this.

However, as I discussed in my July post, for infrastructure & platform services such as Google App Engine for Business, scheduled maintenance still remains excluded in all SLAs:
It is 2011.  If we mark the beginning of Cloud Computing by the initial public release of EC2 (2006), I think enough time has passed for Cloud service providers to do a better job of managing planned outages in a non-service-disruptive way. 

--------
Sidebar - I am a regular user of AWS, GAE, Force.com… I have been using these services for more than a couple of years.  To be fair, I have never received any emails from any of the major cloud for scheduled downtime.  I have received a few from other service providers.  So, I would say that they are all doing a pretty good job operationally (a lot better than probably what most enterprises would do), and make sure the services are always up, and almost always perform well :-].  So, they just have it in the SLA agreements for legal protection & liability.   

Never-the-less, when it comes to migrating or designing enterprise solutions, depending on the application type and use-case, this can become an issue, and require both technical implementation & operations planning.

Sunday, October 17, 2010

A brief look at Oracle and its Cloud Strategy

For the last couple of years, Oracle has shown a consistent strategy to Cloud Computing.  It has made strategic acquisitions such as Virtual Iron to gain x86 virtualization management software, and has also made investments in new products such as Virtual Assembly Builder to facilitate configuration and governance of virtual environments.

Oracle has made clear that it intends to be a provider of technology to both enterprise customers and service providers.  That it does not plan to be a public cloud provider/operator like AWS or Savvis.  Instead, Oracle works with public cloud services as a distribution and delivery partner. 

[Note: See AWS/Oracle announcement of  support for Oracle middleware and apps on EC2 using Oracle VM images.]

---------

This post is a brief look at Oracle and its cloud strategy.  First, I will review Oracle business, financial, and what it brings to Cloud Computing.  Next, I will provide a 5-minute SWOT analysis of Oracle Cloud Strategy.

Oracle business

Oracle’s goal is to be the world’s most complete, open and integrated enterprise software and hardware company.  In FY2010, it booked more than $26B in revenue.  The company breaks down its revenue as follows:

  • Software
    • New software sales
    • Software license renewals & support contract
  • Hardware
    • Hardware sales
    • Hardware support & maintenance contract
  • Services
    • Consulting
    • Education
    • On Demand

Here is their revenue trend for the last 5 years:

image

  • Oracle made over 32% of its 2010 total revenue ($26,821 million) from database & middleware renewal, and about 16% from Fusion apps renewal.  This is due to the fact that almost 90% of Oracle customers renew contracts.   [Renewal has a margin of 85%, and is the key factor to Oracle’s overall profitability.]
  • Oracle’s On Demand, which is where it offers hosted Fusion application, has also contributed to about 3% of total revenue.  This segment has shown steady growth.  In fact, since 2005, it has grown almost 3 times.  [This is a key area to future growth for Oracle especially considering all the investments they have been making to standardize Fusion apps on Fusion middleware and continue to “SaaSify” the applications.]

Oracle Cloud Business Strategy

As stated earlier, Oracle intends to be primarily a Cloud technology provider/enabler as opposed to a service operator.  This was further evidenced as it halted the rollout plans for Project Caroline after the Sun acquisition.

For enterprise customers, Oracle is addressing the needs for private cloud by providing integrated machines such as Exadata and Exalogic.  These machines help customers consolidate workloads and scale up/down as demand grows.  Oracle also continues with new products and enhancement of its middleware and enterprise management solution to enable customers build private clouds on their own hardware.

I think Oracle will be forced to change their cloud strategy for the following reasons:

  • According to analysts, about 10% of IT budget is spent on external cloud services and that percentage will keep growing (see Gartner’s survey).  Oracle needs to pay attention to this shift in enterprise IT spending, if it plans to increase marketshare and revenue.  Large customers struggle with supporting workloads in the cloud, so they look for a vendor to not only help them move workloads to the Cloud, but also provide support and management.  So, they look for hosted managed private clouds.  Oracle could address this need by leveraging Sun assets such as Caroline to offer such services.  This would position it well for future growth.   
  • It is common knowledge that Oracle wants to reach $100B in revenue in the next 10 years.  Cloud Computing, including integrated systems, is a key growth strategy for Oracle.    Oracle needs to diversify to reach that level of revenue in the next decade.  It can’t rely on acquisitions to make that happen.  Let’s assume Oracle acquires CA.  That would only boost Oracle’s revenue by $4B.  

Sidebar: Let’s play the following scenario.  Let’s assume that on average with every Exalogic Oracle charges $1M for hardware and $3M for software. Furthermore, let’s assume a %20 maintenance revenue per box / year.  If Oracle sold 1000 units every year for the next 3 years, they would book a total of $12B in combined new hardware and software + $4B in maintenance.  Everything else constant, by 2014, Oracle’s revenue would grow by $16B to $42B.  Can they do that? 

Oracle Cloud Solution

The following diagram describes Oracle’s cloud solution model:

image

Oracle models its solutions based on different Cloud service offering.   The diagram is pretty self-explanatory.  At the IaaS level, Oracle Sun hardware,  and virtualization technologies (Virtual Iron + Sun).  Oracle offers other capabilities that are not listed in this diagram such as Virtual Desktop Infrastructure (VDI) and Oracle VM Virtual Box.

In the PaaS layer, Oracle uses a combination of virtualization to isolate workloads and management deployment + grid technologies to enable dynamic resources and scaling for applications.

At the top layer, Oracle and non-Oracle apps can be deployed on this platform.  Oracle Fusion apps are optimized for Oracle Fusion middleware. 

Finally, on the right hand side, there is the management layer…The slide is a cut and paste of Richard Sarwal’s presentation at Oracle OpenWorld.  In that presentation, Richard also mentioned that there other capabilities and solution that Oracle will be offering in the next year (i.e. self-service portal, metering & charge-back, etc)

Oracle Cloud SWOT

In terms of integrated systems, Oracle will face competition primarily from IBM CloudBurst and Acadia.  On the middleware side, IBM offers a similar set of offerings based on WebSphere and Tivoli (i.e. WebSphere CloudBurst, WebSphere Virtual Enterprise, Tivoli Cloud Management stack).   Oracle will face competition from VMware vFabric.   IBM has embraced VMware as a virtualization partner (on x86)whereas Oracle decided to acquire its own virtualization.  That has been a source of friction between the two vendors.

In terms of deployment and support, IBM offers more choices than Oracle:

  • IBM & Oracle both offer enterprise-owned cloud
  • IBM offers managed private cloud services (using customers assets), but Oracle does not.  A customer would have to get a managed services contract from an Oracle partner like Wipro.
  • IBM offers IBM-hosted private cloud, but Oracle does not.  A customer would have to sign a contract with an Oracle Cloud provider like Savvis that offers both hosting and support services.
  • IBM offers a public cloud where multiple tenants share the same infrastructure. This is useful for certain workloads (i.e. email, public website) and cloud scenarios (development and testing).  Oracle doesn’t offer that.  A customer would have to find a Pay-As-You-Go provider like AWS.

So, here is a quick SWOT of Oracle Cloud:

image

Let me know what you think? 

Do you think Oracle can reach $100B in the next 10 years through an acquisition only strategy?   What other challenges do you see in Oracle’s cloud strategy, and selling its middleware machine into the enterprise?

Saturday, July 31, 2010

Cloud Services & Availability SLA – Should scheduled maintenance be excluded from the measurements?

Availability SLA is an important criterion to enterprise customers when selecting a Cloud service provider.  If the service provider doesn’t offer appropriate SLAs, they often don’t even make it to the “short list”. 

As an example, Amazon’s EC2 availability SLA is 99.95%. Amazon also offers transparency by providing a website that publishes up-to-the-minute status on their services and any related issues.  In the case of AWS, Amazon makes the following exclusions:

Amazon EC2 SLA Exclusions

The Service Commitment does not apply to any unavailability, suspension or termination of Amazon EC2, or any other Amazon EC2 performance issues:

  • (i) that result from Service Suspensions described in Section 7.1 of the AWS Agreement;
  • (ii) caused by factors outside of our reasonable control, including any force majeure event or Internet access or related problems beyond the demarcation point of Amazon EC2;
  • (iii) that result from any actions or inactions of you or any third party;
  • (iv) that result from your equipment, software or other technology and/or third party equipment, software or other technology (other than third party equipment within our direct control);
  • (v) that result from failures of individual instances not attributable to Region Unavailability; or
  • (vi) arising from our suspension and termination of your right to use Amazon EC2 in accordance with the AWS Agreement (collectively, the “Amazon EC2 SLA Exclusions”).

If availability is impacted by factors other than those explicitly listed in this agreement, we may issue a Service Credit considering such factors in our sole discretion.

Here is section 7.1 from AWS agreement:

  • “…suspended for the duration of any unanticipated or unscheduled downtime or unavailability of any portion or all of the Services for any reason, including as a result of power outages, system failures or other interruptions…”
  • we shall also be entitled, without any liability to you, to suspend access to any portion or all of the Services at any time, on a Service-wide basis:
    • (a) for scheduled downtime to permit us to conduct maintenance or make modifications to any Service;
    • (b) in the event of a denial of service attack or other attack on the Service or other event that we determine, in our sole discretion, may create a risk to the applicable Service, to you or to any of our other customers if the Service were not suspended; or
    • (c) in the event that we determine that any Service is prohibited by law or we otherwise determine that it is necessary or prudent to do so for legal or regulatory reasons (collectively, “Service Suspensions”).

To the extent we are able, we will endeavor to provide you email notice of any Service Suspension in accordance with the notice provisions set forth in Section 15 below and to post updates on the AWS Websites regarding resumption of Services following any such suspension, but shall have no liability for the manner in which we may do so or if we fail to do so.

So, if there is any outage unintended or intended by Amazon, those numbers may not be included in their service availability measurements.  In fairness to Amazon, I haven’t experienced much outage or remember any notices for scheduled maintenance.   Never-the-less, if it occurs, it won’t be counted as an outage.

You can find EC2's SLA at http://aws.amazon.com/ec2-sla/.
--------------------------------
The other day, I received the following notification from GoGrid:
image
In this case, customers basically had no administration access to their running servers for 4 hours. 

GoGrid claims to offers 100% uptime in their SLA, but there are some exclusions as follows:

  1. downtime during scheduled maintenance or Emergency Maintenance
  2. outages caused by acts or omissions of Customer, including its applications, equipment, or facilities, or by any use or user of the Service authorized by Customer
  3. outages caused by hackers, sabotage, viruses, worms, or other third party wrongful actions
  4. DNS issues outside of GoGrid's control
  5. outages resulting from Internet anomalies outside of GoGrid's control
  6. outages resulting from fires, explosions, or force majeure
  7. outages to the Customer Portal
  8. failures during a "beta" period

According to item #1, scheduled maintenance is not included as part of their availability SLA.

Rackspace also claims to offer 100% availability in their SLA, but they also exclude scheduled maintenance from their availability measurement.

So, I appreciate the difference between scheduled maintenance and an unexpected outage.  However, from an availability perspective, I think scheduled maintenance should be included in the availability measurements and reporting.  Otherwise, it is misleading clients.

The Cloud service provider has options to manage for continuous availability.  That is under their control.  Regardless of scheduled vs. unscheduled service outage, the business impact may be no less even if the subscriber is aware of an outage in advance.

What is your experience with Cloud service providers?   Do they explain the SLAs clearly?  Do you think scheduled maintenance should be included or excluded?

Are you considering your application migration options carefully when moving to the Cloud?

So, you’ve heard about the Cloud. You’ve done some prototyping on AWS, Rackspace, GoGrid, Joyent, GAE, Force.com, Engine Yard,…

You show it to your boss. Bada Bing Bada Boom!

This is very timely, because the boss has just come out of a meeting with IT finance. He’s got a big problem justifying the cost of running the company’s website for $2M / year. He’s also pounded daily by business, because he has been unable to meet the availability and performance SLAs despite spending a lot on the infrastructure... So, you’ve just given him a brilliant idea. He checks with Legal, and asks you to look into moving the website to the Cloud.

As an experienced architect, you lift up the hood to take a good look inside the WebApp. You familiarize yourself with the code, dependencies, packaging, etc… After your analysis, you break down the options as follows:
Options
Pros
Cons
Option 1: Move the application as is Quick Inherits the issues that already exist with the application
The performance of the application may just marginally improve due to the application architecture & implementation
Option 2: Re-factor the application; then move Potentially fixes some of the application issues
Potentially fixes some of the application infrastructure issues
Takes more effort than option 1, and requires more time
May introduce some dependencies on the target Cloud
Will probably require learning a few new things
Option 3: Rewrite à Redesign/rewrite the application Full application refresh: New application & infrastructure architecture & design, and implementation
Removes the implementation constraints that existed in previous options thereby allowing to leverage/offer new capabilities (i.e. social computing)
May be more involved than the previous options (time, cost)

In option1, you settle on an infrastructure as a service (IaaS) provider, and just move the application as-is over. Unfortunately, the issues that had already existed with the application will be all propagated (i.e. poor image loading / bundling, deprecated code, unsupported utility jars, packaging). However, in option1, you’ll be able to reduce the infrastructure and operations costs dramatically, and improve SLAs (i.e. availability, auto-scaling) quickly without a lot of efforts.

In option2, you’d still subscribe to an IaaS. You will have the opportunity to clean up the application a little. You might even re-factor the application to use Cloud services (i.e. persistence layer). If time permits, you might even make some changes in the infrastructure in term of content caching & media delivery (i.e. CDN). With this approach, you will have an opportunity to make quick, incremental improvements to the existing application without spending too much money.

In option3, you might consider an IaaS or a PaaS. In this context, the selection is primarily dependent on the application requirements, control over the infrastructure, execution environment customization, etc

With option3, the approach may range from just rewriting the application using its existing design to complete re-implementation. In the case of just rewrite, it can be very straight forward and relatively quick (using the right frameworks, reusing existing components and graphics, etc). The disadvantage is that you’re constraint by the original design, and will not be able to introduce any enhancements or new capabilities.

With redesign/rewrite approach, you’ll need more time to design the solution, but you’ll be able to introduce new capabilities.

So, here is a diagram to summarize:
image

The Time axis is self-explanatory. It reflects effort and costs. The Value axis is an indicator of business value. That ranges from cost-effective & enhanced IT service delivery to offering new business capabilities and possibilities.

I am leaving some details out, but you get the idea.   Let me know what you think.  Tell me about your experience.