Monitoring & Alerting Management Service-Level Expectation (SLE)

A knowledge base article about Monitoring & Alerting Management Service-Level Expectation (SLE) provided by the UC Berkeley IT Service Hub - Knowledge Portal

General Overview

This document describes what is provided by Berkeley IT (bIT) as part of Monitoring and Alert Management support packages, including maintenance, costs, communication, security, incident response, Customer and Service Provider responsibilities, and other related details.

 

Purpose of Agreement

This Service Level Expectation (SLE) document defines the services and support package service levels provided by the Berkeley IT Observability (ITO) Team to its Customers.  Eligible Customers are the units, departments, and colleges internal to the University of California, Berkeley, or any UC-affiliated campus.  This SLE is designed to cover service related terms and conditions, costs, roles and responsibilities, and provides a framework for communication, problem escalation, and service resolution.

 

Length of Agreement

Support Packages are provided on a month-to-month basis and will continue until either party terminates.  The agreement begins when the requested service is provisioned and will automatically renew each month.  Support Packages may be canceled for the next period with a minimum of two business days advance notice.

 

Billing Cycle

Customers will be billed monthly at the beginning of each month.  Any billing questions should be directed to istbill@berkeley.edu.  Credits and Charges cannot exceed sixty (60) days retroactively as per University policy.

 

Service Description

The IT Observability Service offers enterprise level, standards-based, professionally managed Monitoring and Alerting service offerings for devices, systems, networks, and websites within the UC Berkeley and affiliated campus environments.  Support services include (but are not limited to) the agentless monitoring of hardware, virtual instances, ports, services, processes, websites, networks, and the connectivity between all resources configured to use them.  Alerting and notifications for monitored items are available to individuals or groups through multiple vectors such as email, text/SMS, voice, and mobile app.

Additionally, customized monitoring agents, log and config file monitoring, reporting, dashboards, service and topology mapping are available, as well as advanced communication workflows for large-scale, two-way alerting and notification solutions.  Please refer to the MASS COMMUNICATIONS & EMERGENCY NOTIFICATIONS Service for any critical alerting needs, as emergency notifications this service is not supported.  Any services provided outside the scope of this SLE are subject to an additional fee.  Under normal circumstances, the IT Observability staff is available during business hours.  

 

Rates and Fees

Current recharge rates and fees for any associated service offerings can be located in the System Support & Consultation section of the Information Technology Rates page: https://technology.berkeley.edu/rates

 

Hours of Coverage

 

Key Metrics

IT Observability strives for a goal of a 99.99% uptime. Below are defined priorities and expected response times based on Impact and Urgency:

 

Priority Assignment

 

Matrix of choices aligning priority of issue to impact and urgency

Priority

Impact

1 - High

2 - Medium

3 - Low

Urgency

1 - High

1 - Critical

2 - High

3 - Moderate

2 - Med

2 - High

3 - Moderate

4 - Low

3 - Low

3 - Moderate

4 - Low

5 - Planning

 

Service Request and Incident Communications

Service Request is any request made by a Customer to the IT Observability team for routine operational support. 

An Incident refers to an outage or degradation of the normal function of the server or service where it is severely malfunctioning. 

To make a Service Request or report an Incident, the Customer must create a ticket in the ServiceNow ticketing system and provide information that the request is for a Managed Server Support customer, and include the server name (if known).  Any of the following methods are appropriate for submitting a Service Request or reporting an Incident:

  1. Create a ticket by sending an email:
  2. Create a ticket by submitting through the ServiceNow portal:
  3. Create a ticket by contacting the IT Service Desk by phone:
    • Call 510-664-9000
    • Ask to assign this request to “IT Monitoring” or “IT Notifications” as appropriate, or follow prompts for after-hours support.
    • NOTE: It is advised to call the Service Desk directly when creating "Priority 1" or "Priority 2" incidents, as this will provide the quickest response.


If an Extended system support package has been purchased, the Service Desk can be contacted outside of regular business hours for Incidents only.  Be sure to provide the customer name, department, specification that this is for a Extended Support customer, and include the server name (if known).  This information will allow the operators to properly escalate the call to the appropriate support team.

During normal business hours, Service Requests will be responded to within four (4) hours after initial notification to the Service Provider.  Service Request changes will be performed during normal business hours.  Requests placed after normal business hours may not be responded to until the following business day.  If an Incident or Service Request is not responded to with the response times outlined above, the Customer may escalate by directly contacting their assigned ITO technical/functional contact, ITO account manager, or the ITCS Help Desk.  Please refer to the ticket number when escalating.

Incidents reported after normal business hours may not be addressed until the following business day.  After-hours requests for support and emergency support will be fulfilled on a best-effort basis.

Should there be a dispute about services rendered, escalations can be requested by the Customer directly to the management of IT Observability Services.

 

Customer Communications

Depending on the request, one or more of the following methods may be used for communication:

Customers are responsible for providing updated contact information when technical or functional owners change.  Failure to do so may cause delays in service.


Routine Maintenance

Because the central IT environment is regularly upgraded to allow for growth and change in the use of information technology, the Customer should expect routine maintenance to be scheduled periodically in order to comply with new standards and upgrades.  bIT will notify the Customer when such work is needed.  Growth or change initiated by the Customer may warrant a Service Review of their current environment.  Please make note that maintenance work may cause service disruptions. Service outages are published to the bIT system status page at Status Dashboard.


Data Backups

By default, monitoring data points are held by the SaaS provider for a 2-year period, losing granularity and becoming more generalized with age.  Specific configuration data for all devices being monitored in the environment are backed up by the SaaS provider, but with no guaranteed recovery in case of catastrophic failure.  It is therefore highly encouraged to keep all custom configurations and settings noted in your records.

Likewise, specific configuration data for all users, communication methods, and connection strings being used in the notifications environment are backed up by the SaaS provider, but with no guaranteed recovery in case of catastrophic failure.  It is therefore highly encouraged to keep all custom configurations and settings noted in your records.

 

Business Continuity and Disaster Recovery

By default, IT Observability does not provide a solution for Business Continuity or Disaster Recovery solutions with these service offerings.  For further clarification, please refer to the IS-12 requirements associated with the specific application in question.


Security Standards

IT Observability refers to the Minimum Security Standards for Electronic Information (MSSEI) when assigning Security Controls on connected Systems.  Please see the bIT Policy website for more information on minimum security standards (Information Security Office).

As defined in the Roles and Responsibilities Policy, each Customer is the IT Resource Proprietor, and is expected to use their professional judgment in managing risks to the information and systems they use and/or support in partnership with the relevant Institutional Information Proprietor(s). As a Service Provider for the systems, IT Operations will make recommendations based on the customer-identified data classification and risk assessment.  IT Observability will work with the Information Security Office to implement Customer approved solutions to protect the data.

 

Responsibilities and Expectations

IT Observability:

  1. Includes one hour of consultation to gather requirements and provide a cost estimate for the initial service engagement.
  2. Offers server provisioning and operating system management.
  3. Operating system patches are reviewed for their criticality as they are released.
    • Routine patches are applied on a consistent basis in order to minimize server outages.
    • Security patches deemed critical may be applied outside the predefined maintenance window or patching schedule.
    • bIT Platform Infrastructure will not be responsible for any application failure, downtime or issues resulting from these mandatory updates.
  4. Provides Monitoring and Alerting focused technical support and problem resolution.
  5. Provides no management or system support services for any monitored systems.
  6. Maintains server security in accordance with the policies governing UC Berkeley Information Technology resources.
  7. Does not assist with server lifecycle management. 
  8. Provides a comprehensive end-to-end monitoring and a notification tools for system performance, network devices, services & processes, and connectivity.  All monitoring can be configured for alerting, but there is no implied or explicit support for the monitored systems by any part of the IT Observability staff.
  9. Is not responsible for the administration or configuration of any bSecure network firewall systems affecting connectivity between data collectors and monitored systems.
  10. Is not responsible for any network Load Balancer configurations affecting monitoring efforts.
  11. Coordinates with other bIT Departments and 3rd party vendors as needed.
  12. Provides best effort system and/or application monitoring & notification configurations.  If further assistance is required, ITO will refer the customer to the appropriate vendor for additional assistance.
  13. Perform planned maintenance on a scheduled basis for each managed portal as per the schedule of the SaaS vendor.  Any pending maintenance activities will be available to the customer through the vendor website, or upon login to the portal.
  14. Provide the Customer contact with notification of any service disruptions and emergency maintenance as soon as feasible.
  15. User accounts are only provided to individuals with active UC Berkeley Affiliations and CalNet accounts.

In order to protect the interests and assets of the University of California Berkeley, bIT may be required to render services beyond those described in this document. Such additional support is provided at the discretion of University senior management with Customer consultation. This work may result in additional charges.

The Customer:

  1. Acts as the IT Resource Proprietor of their data, and must categorize that data according to University requirements while working with the Institutional Information Proprietor(s).
  2. Is responsible for the installation, configuration, maintenance, patching, upgrades, licensing, and security of their installed applications.  Any assistance from bIT required to meet these obligations may involve a Time and Materials (hourly) charge.
  3. Is required to perform application testing for all patches, upgrades, and database changes in a timely manner.
  4. Is responsible for notifying their own application users and/or Service Desk of any service interruptions or outages, as appropriate.  
  5. Is required to provide up to date contact information. 24x7 support requires 24x7 contacts.
  6. Is responsible for approving access to the system resources and application data.
  7. Is responsible for promptly notifying ITO when user or service account access or permissions should be changed due to UC Berkeley Affiliation expiration or changes. 
  8. Is responsible for providing a Technical/Functional Contact to assist in troubleshooting issues.
  9. Is responsible for providing an Information Security Contact and responding to ISO alerts with regard to their application.
  10. Is responsible for providing a designated Billing Contact.
  11. The Customer must register applications hosting restricted data in the Socreg Asset Registration Portal.
  12. Is responsible for complying with all campus computer use and security policies.  bIT reserves the right to shut down or isolate any server that is found to be out of date, vulnerable, or compromised.
  13. Will provide prompt payment and provisioning of appropriate chartstring (please note: chartstrings or “COAs” are required to request provisioning of certain services.) 
  14. Will respond to IT Observability staff inquiries in a timely manner.
  15. Will consult with IT Observability before making hardware or software purchases related to supported systems.
  16. Should anticipate Monitoring or Notifications vendor changes/renewals every 3-5 years. 
  17. Is solely responsible for submitting security exceptions to ISO if ITO staff makes the customer aware of vulnerabilities during routine support efforts.