Begin with a named inventory and an onboarding review
List each server, environment, operating system, database and supported application. Include domains, certificates, backups and external dependencies where they are in scope. Identify unsupported versions, undocumented custom changes and missing access before accepting responsibility. Agree which inherited problems will be corrected, priced separately or explicitly excluded. The inventory should say who owns each account and where operating information is maintained. A general promise to manage the server leaves too much room for disagreement.
Define routine work as a repeatable process
Specify which updates are applied, how their urgency is assessed and who checks compatibility. State the maintenance window, notice process and conditions requiring a restart or customer approval. Include certificate renewal, capacity checks, log handling and access review where needed. For a failed update, identify the rollback route and who tests the application afterward. Make a distinction between maintaining the current supported system and a major migration that changes its architecture or software generation.
Separate detection, acknowledgement and restoration
Define what monitoring covers and which events create an alert. Explain who is available, how an incident is prioritized and which channel reaches that person. A response target measures when someone starts handling the issue; it should not be presented as a promise that every cause will be fixed by then. Record the escalation route to the infrastructure provider and application team. Google’s incident guidance is useful context for clear roles and communication, but your coverage must be stated in your own agreement.
Write a recovery section that can be rehearsed
Name the data, configuration and credentials included in backups, their location and the retention policy. Define the recovery point objective as the acceptable amount of lost recent data, and the recovery time objective as the target time to restore the agreed service. Ask which failure scenario the targets cover and what test supports them. Specify who authorizes restoration and how recovered data is verified. A successful backup job alone does not establish that the complete application can be recovered.
Plan a relevant restore exercise
Use backup and recovery work to check the agreement’s data-loss and interruption assumptions.
Explore Plan a relevant restore exerciseExplain how requests and costs are controlled
Separate routine maintenance, incident work and planned improvements. State how urgent work changes the queue, whether available effort is capped and what happens when it is exhausted. Clarify approval for extra capacity, licenses or specialist help. Require a short record of material configuration changes and the reason for them. That record should be useful to the next operator as well as the customer. Avoid a scope where a problem is monitored but every action needed to resolve it is undefined.
Make acceptance and exit practical
Before the service starts, verify the inventory, access route, alert delivery and a relevant recovery procedure. Keep unresolved onboarding issues visible with an owner and next action. For exit, define the delivery of configuration, runbooks, backup access and account information, followed by access revocation. The business should be able to identify what runs where and who can operate it. Recheck the agreement when a new application or provider is introduced so responsibilities remain current.
Explore server management
Use a concrete operating scope to discuss ongoing server responsibility.
Explore Explore server managementCompare hosting arrangements
Decide which account and provider arrangement should sit behind the management agreement.
Explore Compare hosting arrangements