A knowledge base article about Monitoring & Alerting Management Service-Level Expectation (SLE) provided by the UC Berkeley IT Service Hub - Knowledge Portal
This document describes what is provided by Berkeley IT (bIT) as part of Monitoring and Alert Management support packages, including maintenance, costs, communication, security, incident response, Customer and Service Provider responsibilities, and other related details.
This Service Level Expectation (SLE) document defines the services and support package service levels provided by the Berkeley IT Observability (ITO) Team to its Customers. Eligible Customers are the units, departments, and colleges internal to the University of California, Berkeley, or any UC-affiliated campus. This SLE is designed to cover service related terms and conditions, costs, roles and responsibilities, and provides a framework for communication, problem escalation, and service resolution.
Support Packages are provided on a month-to-month basis and will continue until either party terminates. The agreement begins when the requested service is provisioned and will automatically renew each month. Support Packages may be canceled for the next period with a minimum of two business days advance notice.
Customers will be billed monthly at the beginning of each month. Any billing questions should be directed to istbill@berkeley.edu. Credits and Charges cannot exceed sixty (60) days retroactively as per University policy.
The IT Observability Service offers enterprise level, standards-based, professionally managed Monitoring and Alerting service offerings for devices, systems, networks, and websites within the UC Berkeley and affiliated campus environments. Support services include (but are not limited to) the agentless monitoring of hardware, virtual instances, ports, services, processes, websites, networks, and the connectivity between all resources configured to use them. Alerting and notifications for monitored items are available to individuals or groups through multiple vectors such as email, text/SMS, voice, and mobile app.
Additionally, customized monitoring agents, log and config file monitoring, reporting, dashboards, service and topology mapping are available, as well as advanced communication workflows for large-scale, two-way alerting and notification solutions. Please refer to the MASS COMMUNICATIONS & EMERGENCY NOTIFICATIONS Service for any critical alerting needs, as emergency notifications this service is not supported. Any services provided outside the scope of this SLE are subject to an additional fee. Under normal circumstances, the IT Observability staff is available during business hours.
Current recharge rates and fees for any associated service offerings can be located in the System Support & Consultation section of the Information Technology Rates page: https://technology.berkeley.edu/rates
IT Observability strives for a goal of a 99.99% uptime. Below are defined priorities and expected response times based on Impact and Urgency:
Priority |
Impact |
|||
|---|---|---|---|---|
|
1 - High |
2 - Medium |
3 - Low |
||
Urgency |
1 - High |
1 - Critical |
2 - High |
3 - Moderate |
|
2 - Med |
2 - High |
3 - Moderate |
4 - Low |
|
|
3 - Low |
3 - Moderate |
4 - Low |
5 - Planning |
|
A Service Request is any request made by a Customer to the IT Observability team for routine operational support.
An Incident refers to an outage or degradation of the normal function of the server or service where it is severely malfunctioning.
To make a Service Request or report an Incident, the Customer must create a ticket in the ServiceNow ticketing system and provide information that the request is for a Managed Server Support customer, and include the server name (if known). Any of the following methods are appropriate for submitting a Service Request or reporting an Incident:
If an Extended system support package has been purchased, the Service Desk can be contacted outside of regular business hours for Incidents only. Be sure to provide the customer name, department, specification that this is for a Extended Support customer, and include the server name (if known). This information will allow the operators to properly escalate the call to the appropriate support team.
During normal business hours, Service Requests will be responded to within four (4) hours after initial notification to the Service Provider. Service Request changes will be performed during normal business hours. Requests placed after normal business hours may not be responded to until the following business day. If an Incident or Service Request is not responded to with the response times outlined above, the Customer may escalate by directly contacting their assigned ITO technical/functional contact, ITO account manager, or the ITCS Help Desk. Please refer to the ticket number when escalating.
Incidents reported after normal business hours may not be addressed until the following business day. After-hours requests for support and emergency support will be fulfilled on a best-effort basis.
Should there be a dispute about services rendered, escalations can be requested by the Customer directly to the management of IT Observability Services.
Depending on the request, one or more of the following methods may be used for communication:
Customers are responsible for providing updated contact information when technical or functional owners change. Failure to do so may cause delays in service.
Because the central IT environment is regularly upgraded to allow for growth and change in the use of information technology, the Customer should expect routine maintenance to be scheduled periodically in order to comply with new standards and upgrades. bIT will notify the Customer when such work is needed. Growth or change initiated by the Customer may warrant a Service Review of their current environment. Please make note that maintenance work may cause service disruptions. Service outages are published to the bIT system status page at Status Dashboard.
By default, monitoring data points are held by the SaaS provider for a 2-year period, losing granularity and becoming more generalized with age. Specific configuration data for all devices being monitored in the environment are backed up by the SaaS provider, but with no guaranteed recovery in case of catastrophic failure. It is therefore highly encouraged to keep all custom configurations and settings noted in your records.
Likewise, specific configuration data for all users, communication methods, and connection strings being used in the notifications environment are backed up by the SaaS provider, but with no guaranteed recovery in case of catastrophic failure. It is therefore highly encouraged to keep all custom configurations and settings noted in your records.
By default, IT Observability does not provide a solution for Business Continuity or Disaster Recovery solutions with these service offerings. For further clarification, please refer to the IS-12 requirements associated with the specific application in question.
IT Observability refers to the Minimum Security Standards for Electronic Information (MSSEI) when assigning Security Controls on connected Systems. Please see the bIT Policy website for more information on minimum security standards (Information Security Office).
As defined in the Roles and Responsibilities Policy, each Customer is the IT Resource Proprietor, and is expected to use their professional judgment in managing risks to the information and systems they use and/or support in partnership with the relevant Institutional Information Proprietor(s). As a Service Provider for the systems, IT Operations will make recommendations based on the customer-identified data classification and risk assessment. IT Observability will work with the Information Security Office to implement Customer approved solutions to protect the data.
In order to protect the interests and assets of the University of California Berkeley, bIT may be required to render services beyond those described in this document. Such additional support is provided at the discretion of University senior management with Customer consultation. This work may result in additional charges.