Cloud Migration Services in India
Discovery, dependency mapping, landing zone design, database and stateful workload moves, rehearsed cutovers and a rollback you have actually tested. We run migrations to AWS, Azure and Google Cloud from India for engineering teams in the US, UK, Canada, Australia and New Zealand.
What Cloud Migration Services in India Actually Buy You
Cloud migration is the work of moving applications, data and the operations around them from where they run now to a cloud platform, without losing data and without a week of unplanned downtime. It covers finding out what you actually have, deciding per application whether to move it as it is or change it on the way, building the account and network structure it will land in, moving the data, switching traffic over, and proving you can go back if the switch goes badly.
If you are reading this, you are probably a CTO or engineering manager somewhere between London and Auckland with a hardware refresh, a data centre lease, a VMware renewal or a board deadline pushing you. You have a rough plan. What you do not have is the four to six weeks of unglamorous discovery work that turns the rough plan into something you can commit a date to, and you do not want to hire three permanent cloud engineers for a programme that ends.
That is the shape of work our teams in India take on. Not advice about the cloud in general. A named group of engineers who inventory the estate, argue with you about which of the six Rs each application deserves, build the landing zone in code, move the workloads in waves, sit through the cutover, and hand you runbooks your own team can operate afterwards.
The section headings below are the order the work actually happens in. If your migration is already underway and stalled, the discovery and cutover sections are usually where the missing piece is.
The Problem Underneath the Project
Almost nobody migrates because they want to. Something forces it. The colocation contract has eighteen months left and renewing means another five year commitment. A storage array is out of vendor support and the quote for the replacement is a capital request nobody wants to write. Virtualisation licensing costs jumped after a vendor change and the finance director has noticed. Or the product team cannot ship because provisioning a test environment takes six weeks and a purchase order.
Then the project starts and hits the same wall every time. Nobody can produce a reliable list of what is running. There is a server everyone calls "the reporting box" that has been up for four years, has no owner, and turns out to be the thing that sends the daily settlement file. There is an application whose vendor went out of business in 2017 and whose licence key is tied to a MAC address. There is a scheduled task on someone's laptop.
The cost of that unknown is not abstract. A migration wave that has to be aborted at 04:00 because an undocumented dependency reached back to an on premise file share burns a weekend, burns credibility with the business, and pushes the next wave out by a month because nobody will approve a Saturday until the cause is understood. Two or three of those and the programme acquires a reputation, and reputational damage inside an engineering organisation is what turns an eight month migration into a two year one.
The second cost arrives later and lands on the finance team. The estate moves, the bill arrives, and it is higher than the data centre it replaced. That is not a cloud problem. It is the entirely predictable outcome of copying on premise sizing into a place that charges by the hour, and it has its own section further down because it deserves one.
The Six Rs, and When Each One Is the Wrong Call
Gartner published five migration options in 2010 and AWS popularised a six strategy version in 2016, which is where most of the vocabulary in your steering committee slides comes from. The framework is useful. What is missing from most presentations of it is the other half: each strategy has conditions under which it is actively the wrong answer, and choosing wrongly is more expensive than choosing slowly.
Rehost, or lift and shift
Move the virtual machine as it is. Same operating system, same application binaries, same configuration, different hypervisor underneath. Block level replication tools do most of the work, and a well drilled team can move dozens of servers a wave.
Right when: a lease or support deadline is fixed and immovable, the application is stable and nobody is developing it, or you need to prove the migration muscle works before attempting anything harder. It is also the correct answer for the long tail of small applications where the engineering time to improve them will never be repaid.
Wrong when the application was already scheduled for a rewrite, because you will pay to move it and then pay again to replace it. Wrong for anything whose cost driver is architectural, such as a monolith that needs sixteen cores to serve two hundred users, since rehosting preserves the shape and therefore the bill. And wrong where licensing terms change badly on shared tenancy hardware, which is a real constraint with some database and middleware vendors and needs checking against your specific agreement before you commit.
Replatform, sometimes called lift and reshape
Move it, but swap one or two components for managed equivalents on the way. A self managed MySQL becomes RDS or Cloud SQL. A hand built load balancer becomes the platform's. The application code stays largely untouched.
This is the best value per hour of effort in most migrations. The operational burden on a typical estate is concentrated in a handful of undifferentiated components, and handing patching, backup and failover of a database to the platform removes more toil than anything else you could do with the same week.
It goes wrong when the managed service quietly does not support something you depend on. Managed database services restrict superuser access, plugin installation and some replication configurations. If your application relies on a specific extension, a custom collation, a linked server or a scheduler running inside the database, find that out during discovery, not on cutover night. The other trap is a swap sold as low risk that also changes the database engine version, turning one variable into two.
Refactor or re-architect
Change the application to suit the platform. Break out services, move to containers or serverless, replace a queue, restructure the data layer.
Right when the application is strategic, actively developed, and the current architecture is what limits throughput or release frequency. Also right when the target state genuinely needs elasticity that the current design cannot express.
Wrong as part of the migration itself for anything you cannot afford to have in flight for two quarters. The single most common way migration programmes lose their deadline is combining a data centre exit with a rewrite, so that a fixed external date now depends on an open ended engineering effort. Move it first, change it second, and let the two have separate deadlines. Wrong also where the team that will own the refactored system was not part of designing it, because you will have built something nobody wants to operate.
Repurchase
Stop running it and buy the software as a service equivalent. Self hosted mail, wikis, ticketing, HR systems and some CRM installations all fall here.
If the application is not a differentiator and a mature product exists, buy it. The migration cost is mostly data extraction and user retraining rather than engineering, and you stop owning a system that was never your business.
The standard trap is a long lived ticketing or ERP installation carrying a decade of accumulated configuration that quietly encodes real business rules. Nobody can list those rules, so nobody can specify the replacement, and the project turns into archaeology. There is a second constraint worth checking early: whether data residency or retention obligations conflict with where the vendor stores data. That question belongs with your counsel and your data protection officer rather than with the migration team, and it needs asking before the contract, not after.
Retire
Turn it off. On most estates somewhere between ten and twenty percent of servers are running something nobody uses any more, and discovery is what surfaces them.
The cheapest workload to migrate is the one you delete. Anything with no inbound connections across a full business cycle and no owner willing to claim it is a candidate.
The danger is measuring "no traffic" over a fortnight when the thing runs quarterly. So we watch for at least a month, longer where a financial or regulatory cycle is involved, and we power off before we delete, keeping the storage for an agreed period so the decision stays reversible. Where a system holds records you are legally obliged to retain, the answer is an archive with a retention policy rather than a deletion, and that distinction is worth confirming with whoever owns your retention schedule.
Retain
Leave it where it is, deliberately, and write down why and when the decision gets revisited.
Right for workloads with hardware dependencies, sub millisecond latency requirements to on premise equipment, licence terms that make cloud hosting uneconomic, or a compliance position not yet resolved. Also right for anything being decommissioned within the year.
Wrong when it is used to avoid a difficult conversation. A retained estate of eleven servers means the data centre stays open, the network link stays, the support contract stays, and the business case for the whole programme quietly evaporates. If retaining anything means keeping the facility, the retain decision has to be argued against that full cost, not against the cost of moving that one application.
You will also see a seventh R, relocate, for moving whole VMware estates onto a cloud hosted equivalent without converting the machines. It is a legitimate route for a fast data centre exit, and it carries the same warning as rehosting: it is a starting position, not a destination.
Discovery and Dependency Mapping
This is the phase that gets cut when a deadline is tight, and cutting it is the reason the deadline is then missed. Discovery produces three artefacts: an inventory of what runs, a dependency graph of what talks to what, and an application grouping that tells you what has to move together in the same wave.
Inventory, and why the CMDB is not it
Most organisations have a configuration management database and most of them are wrong, because they record what was provisioned rather than what is running. We collect from the hypervisor and the network in parallel: vCenter exports, cloud provider APIs where part of the estate has already moved, DHCP and DNS records, switch ARP tables, and agent based collectors where they are permitted.
The tools that do this properly are AWS Application Discovery Service with its agentless collector, Azure Migrate's appliance based discovery, Google's StratoZone assessment, and vendor neutral options such as Device42 or Flexera where the target platform has not been chosen yet. The vendor neutral route is worth the extra cost specifically when you are still deciding between AWS and Azure, because the provider tools are built to produce a business case for their own platform and their sizing recommendations reflect that.
Dependency mapping, agent based and agentless
An inventory tells you a server exists. It does not tell you that the payment service opens a connection to a database on a server in a different rack every ninety seconds. Agentless dependency discovery infers relationships by polling connection tables over time, and it is safe to deploy but blind to anything infrequent. Agent based collection, using the Dependency agent with Azure Migrate or equivalent collectors elsewhere, captures process level detail including which binary owns the socket, at the cost of a change management process to install agents on production servers.
Run it for a minimum of four weeks. Two weeks misses month end. On estates with quarterly regulatory reporting we argue for a full quarter of observation, and where that is impossible we compensate with interviews and scheduled job audits, and we say plainly which parts of the map are inferred rather than observed.
Grouping applications into waves
The dependency graph is then cut into move groups. Anything with a chatty, latency sensitive relationship goes in the same wave, because splitting it across the network link between your data centre and the cloud adds round trip time to every one of those calls. A pair of systems exchanging a thousand queries per page render will work acceptably at 0.2 milliseconds apart and fall over at 30.
Wave one should be a real application, not a test server. It should be low risk enough that a bad night is survivable and real enough that it exercises the landing zone, the network path, the monitoring, the runbook format and the cutover process. The point of wave one is not the workload. It is finding out what your migration process gets wrong while the stakes are low.
What discovery hands to you
A server inventory with utilisation data and an owner per system, a dependency map with confidence noted per edge, a per application R decision with the reasoning recorded, a wave plan with sequencing, a target sizing based on observed utilisation rather than provisioned capacity, and a cost model. That last one is what your finance team will hold you to, so it names its assumptions about data transfer, storage tiers and commitment discounts explicitly rather than burying them.
Landing Zone Design: Get These Right Before Anything Moves
A landing zone is the pre built environment your workloads arrive into. The reason it matters is asymmetry: some decisions are trivial to change later and some are close to impossible, and the impossible ones all have to be made before the first server moves.
Account and subscription structure
The unit of blast radius, billing separation and policy attachment is the AWS account, the Azure subscription or the Google Cloud project. Splitting later means moving resources between them, which for many services means recreating them.
The pattern that survives contact with reality is separation by environment and by ownership: production isolated from non production at the account boundary, shared services such as networking and logging in their own accounts, and a security account nobody deploys into. AWS Control Tower and the Landing Zone Accelerator implement this; on Azure it is management groups and the Cloud Adoption Framework landing zone reference; on Google Cloud it is the organisation, folder and project hierarchy with organisation policy constraints. Whichever you use, build it with Terraform or the platform's native infrastructure as code rather than by hand, because you will need to explain in an audit why a control exists and a console click leaves no such record.
Network topology and address space
Pick address ranges that do not collide with your on premise network, with any network you might acquire, and with the ranges your partners use for their VPN endpoints. Overlapping CIDR blocks are the single most tedious problem to fix after the fact, because the fix is renumbering a live environment.
Hub and spoke is the usual topology, with a transit hub carrying the connection back to your data centre. AWS Transit Gateway with Direct Connect, Azure Virtual WAN with ExpressRoute, or Cloud Interconnect on Google. Order the circuit early. Physical cross connects have lead times measured in weeks and they are frequently the thing that delays wave one, particularly for a data centre in a location where the provider has no existing presence.
Identity, before anyone gets an access key
Federate to your existing identity provider rather than creating a parallel set of cloud users. Entra ID, Okta or your Google Workspace directory becomes the source of truth, with roles assumed through single sign on and short lived credentials. Break glass accounts exist, have hardware second factors, and their use pages someone.
This is also where offshore access gets designed. Our engineers work under named identities in your directory with scoped roles, so every action is attributable in CloudTrail or the Azure activity log. Not shared accounts and not a static access key in a chat message.
Guardrails, logging and tagging
Preventive controls stop the mistake happening. Service control policies or Azure Policy denying public storage buckets, denying resource creation outside approved regions, and denying the deletion of log destinations. Detective controls catch what preventive ones miss, through Security Hub, Defender for Cloud or Security Command Center.
Logs go to an account the workload teams cannot write to, with retention set once at the start. And the tagging standard gets agreed before migration rather than retrofitted: owner, environment, application, cost centre, and a data classification tag. Enforce it at creation with policy, because a tag that is optional is a tag that is missing on exactly the resources you later need to attribute.
Moving Data and Stateful Workloads
Stateless servers are the easy part. The migration is really about the state: databases, file shares, message queues mid flight, and the several terabytes of documents that a scanning department has been writing to a NAS since 2011.
Databases, homogeneous
Same engine on both sides is the straightforward case, and it is still not trivial. The mechanism is a full load followed by continuous change data capture until the replica is close enough to the source to cut over. AWS DMS, Azure Database Migration Service and Google Cloud's Database Migration Service all work this way, as does native replication: PostgreSQL logical replication, MySQL with GTID based replication, SQL Server transactional replication or log shipping, Oracle Data Guard or GoldenGate.
Native replication is usually more faithful and the managed tools are usually faster to set up. Where the estate is large and uniform we lean on the managed service; where a single database is business critical and complex we prefer native mechanisms because we can reason about exactly what they do and do not carry.
The prerequisites are where projects lose a week. Change data capture needs supplemental logging enabled on Oracle, binary logging in row format on MySQL, and a replication slot and appropriate WAL level on PostgreSQL. Tables without a primary key will not replicate updates correctly in most tools. And every one of these tools carries data and leaves behind some combination of secondary indexes, foreign keys, triggers, stored procedures, sequences and permissions, depending on configuration. Read the tool's own list of what it excludes, script the remainder, and verify object counts on both sides before you plan a cutover.
Databases, heterogeneous
Changing engine at the same time as changing location is a different class of project. Oracle to PostgreSQL, SQL Server to Aurora, DB2 to anything. The data movement is the smallest part.
The AWS Schema Conversion Tool and its DMS Schema Conversion successor will convert a large share of the schema automatically and produce a report of what it cannot. Open source ora2pg does a comparable job for Oracle to PostgreSQL and is worth running alongside for a second opinion. Babelfish for Aurora PostgreSQL accepts the SQL Server wire protocol and T-SQL, which can remove a large amount of application rewriting where the compatibility coverage suits your code.
The manual residue is always the same categories: PL/SQL or T-SQL packages with vendor specific behaviour, hierarchical queries, autonomous transactions, and application code that depends on the old engine's implicit type conversion, null sorting order or date handling. Budget the application change as the main line item and the data move as a supporting task, not the other way round. A realistic heterogeneous database migration runs nine to eighteen months for a substantial system, and we would rather say so at the proposal stage than discover it together in month five.
File shares, object storage and bulk transfer
For file data, AWS DataSync, Azure's Storage Migration Service and AzCopy, and Google's Storage Transfer Service all handle incremental sync over the network with verification. Robocopy and rsync still have their place for smaller sets and give you complete control over the retry behaviour.
Do the arithmetic before choosing the network. A hundred terabytes over a one gigabit link, achieving perhaps sixty percent of line rate in practice, is around fifteen days of continuous transfer, and that link is also carrying your production traffic. Physical transfer with AWS Snowball, Azure Data Box or Google Transfer Appliance stops being exotic and starts being obvious somewhere around the fifty to a hundred terabyte mark, depending on your bandwidth. Whichever route, plan a final incremental sync immediately before cutover, because the data changes while the bulk copy is in flight.
Preserve what matters: NTFS or POSIX permissions, timestamps, symlinks and alternate data streams. Verify with checksums rather than file counts. A transfer that reports success and quietly dropped access control lists produces a support queue you will be answering for a month.
The things people forget are stateful
Message queues holding undelivered messages at the moment of cutover. Scheduled jobs on servers not in scope. Application servers writing session state to local disk. Hardcoded IP addresses in configuration files, firewall rules and partner allow lists. Certificates pinned to hostnames. Licence keys tied to hardware identifiers. Every one of these is discovered during a properly run dry run and none of them are discovered by reading documentation.
Cutover Strategy: Big Bang, Phased or Strangler
The cutover is the twenty minutes the whole programme is judged on. Three approaches, and the right one is decided by your coupling and your tolerance for a bad night, not by preference.
Big bang
Everything moves in one window. Freeze, final sync, switch, verify, unfreeze.
It suits tightly coupled systems that genuinely cannot be split, and small estates where the window is achievable. Its virtue is that there is no period of running two environments in parallel, which is itself a source of bugs and cost. Its vice is obvious: all the risk lands on one night, and if you exceed the window you are choosing between rolling back and running over into a working day.
If you take this route, build the runbook as a timed script with named owners and a hard go or no go checkpoint partway through, at the point where continuing commits you. Rehearse the whole thing at least twice against a copy of production, and time each step so the window is based on measurement rather than optimism.
Phased by application or by wave
The default for most estates. Move a move group at a time over weeks or months, each with its own window and its own rollback.
The trade off is a hybrid period where half the estate is in the cloud and half is not, which means the network link between them is now on the critical path for production traffic, latency sensitive pairs must not be split, and identity, DNS and monitoring have to work across both. That hybrid period is real work and it is where phased migrations overrun, so it belongs in the plan as a cost rather than as a free consequence.
What you get for it is the ability to learn. Wave three is always smoother than wave one, and the accumulated corrections to the runbook are worth more than any amount of upfront planning.
Strangler, for applications rather than servers
Martin Fowler's strangler fig pattern: put a facade in front of the old system, move one capability at a time behind it, and shrink the original until it can be switched off. In a migration context, a routing layer sends some paths to the new environment and the rest to the old.
This is the right approach when you are refactoring as you migrate and the application is too large to move atomically, and when you need to demonstrate progress to a business that does not want to hear "nothing visible for six months".
It is the wrong approach when the data cannot be split. Two systems writing to what is logically one dataset means dual write, and dual write means you own consistency and conflict resolution yourself. That is a genuinely hard distributed systems problem and it is often harder than the migration you were trying to simplify. If the strangler route requires dual write to a shared database, be certain that is easier than a clean cutover before committing to it.
The mechanics of the switch
Lower DNS time to live to sixty seconds several days ahead, and verify that your clients actually respect it, because some Java runtimes and some connection pools cache resolution far longer than the record says. Drain connections rather than severing them. Have the verification suite written and automated before the night: a set of read checks, a synthetic transaction end to end, row counts and checksums on the critical tables, and a look at the error rate for fifteen minutes before anyone declares success. Keep the source environment running afterwards for an agreed period rather than decommissioning on the Monday.
Rollback: The Plan You Hope Not to Use
Every migration plan has a rollback section. Most of them are one paragraph saying "revert DNS", which is not a plan, it is a hope. Rollback is a design problem and it has to be solved before the cutover, because the moment you need it is the moment you have no time to invent it.
The point of no return
Write down the exact step after which rollback stops being clean. Usually it is the first write accepted by the new system that has not replicated backwards. Before that point, rolling back costs you a night. After it, rolling back costs you a data reconciliation exercise, because the two databases have diverged and someone has to decide which version of the truth wins for every affected record.
Making that boundary explicit changes how the go or no go decision is taken. Before it, the answer to doubt is to roll back. After it, the answer to doubt is to fix forward, and the team needs to know which regime it is operating in without debating it at 04:00.
Reverse replication
To extend the clean rollback window past the first write, configure replication from the new database back to the old one and turn it on at cutover. AWS DMS supports a reverse task; native logical replication can be set up bidirectionally with care around conflicts and loops.
Test it before the night. Reverse replication that has never run is a line in a document, and the failure modes tend to be things like sequence values on the old side now colliding with rows written on the new side, which you find out in ninety seconds of testing and would otherwise find out during an incident.
What has to still exist
The source environment, running, with its data current enough to serve. The old DNS records and load balancer configuration, ready to restore. The firewall rules that were removed. A copy of the pre migration state, tested by restoring it somewhere, because an untested backup is a belief rather than a capability. And an agreed retention period for all of it, written down, so nobody decommissions the source on day three because a decommissioning ticket was already in the queue.
Rehearse the rollback, not just the migration
Dry runs almost always rehearse the happy path. We run at least one dry run where the rollback is executed for real, timed, with the same people who will be on the actual night. It surfaces things no document does: the credentials for the old load balancer nobody has, the monitoring that was already pointed at the new environment, the alert suppression window that expired. Rollback is a skill, and the first time anyone performs it should not be under pressure.
Migration Tooling and the Trade-offs
Tool choice matters less than most vendors suggest, and more than most engineers admit. Here is our honest read on what to use when.
AWS Application Migration Service, and the CloudEndure question
AWS Application Migration Service, usually written MGN, is the successor to CloudEndure Migration, which AWS acquired and has since folded into its own service. If a consultant is proposing CloudEndure Migration for a new AWS migration today, that is a signal their playbook has not been updated; MGN is the current path, and CloudEndure Disaster Recovery similarly became AWS Elastic Disaster Recovery.
MGN does continuous block level replication from a source server, agent based, into a staging area, then launches test instances you can boot and validate without touching the source. The strength is that testing is cheap and repeatable, so you can rehearse a wave five times. The limits are that it is a rehost tool and nothing more, agents have to be installed on every source machine, and the staging area has a running cost while replication continues that surprises people who leave a wave replicating for three months.
AWS DMS and Schema Conversion
DMS is the workhorse for database moves into AWS, handling homogeneous and heterogeneous paths, with continuous replication for low downtime cutovers. Its weakness is what it silently does not do: it is deliberately conservative about secondary indexes, constraints, triggers and stored code, and teams that assume "DMS moved the database" without verifying object counts get a nasty surprise on the first slow query. Treat DMS as a data mover and handle schema separately with SCT or ora2pg.
Azure Migrate
Azure Migrate is a hub rather than a single tool: discovery and assessment through an on premise appliance, agentless or agent based dependency analysis, server migration, and an app containerisation path for suitable web applications. Database work goes through Azure Database Migration Service, with Data Migration Assistant for the pre migration compatibility assessment against SQL Server.
It is the most integrated option if you are already a Microsoft estate with Entra ID, System Center and SQL Server, and the assessment quality for Windows workloads is genuinely good. The trade off is the same as with any provider tool: its recommendations point at Azure, and its sizing is generous. Sanity check the output against your own utilisation data before signing off a business case on it.
Google Cloud Migrate to VMs and Migrate to Containers
Migrate to Virtual Machines, previously Migrate for Compute Engine and before that Velostrata, streams a machine into Compute Engine and can start it running before the full copy has completed, which shortens the window meaningfully. Migrate to Containers takes a suitable virtual machine and produces container artefacts and Kubernetes manifests, which is a genuinely distinctive capability where you want to land on GKE rather than on virtual machines.
Assessment runs through StratoZone. Database work uses Google's Database Migration Service, which is strong for MySQL and PostgreSQL and has an Oracle to PostgreSQL path worth evaluating alongside the AWS equivalent if the target platform is still open.
Vendor neutral and third party
Where the platform decision is not yet made, or where the estate spans several providers, Device42 and Flexera give you an inventory and dependency map that is not built to sell you a destination. Carbonite Migrate and Zerto turn up in estates that already own them for disaster recovery. And for VMware heavy estates, the relocate route onto a cloud hosted VMware equivalent avoids machine conversion entirely, at the cost of keeping the licensing relationship you may have been trying to escape.
Our default, stated plainly: use the target provider's native tools for the bulk of the moves because they are free or close to it and well supported, use native database replication for the two or three databases you cannot afford to get wrong, and buy a third party discovery tool only when the platform decision is genuinely open or the estate is too large for the provider's collectors to cover cleanly.
Why Is the First Cloud Bill Always a Shock?
Because the estimate was built from compute prices and the bill is built from everything else. Compute is the easy part to forecast and it is rarely where the overrun lives.
Data transfer, the line item nobody modelled
Traffic leaving a cloud provider to the internet is charged per gigabyte, and that egress charge is the one most business cases omit entirely. It matters enormously for some workloads and not at all for others. A media platform serving video, a backup target receiving restores, an analytics estate shipping extracts to a partner: all of these can spend more on transfer than on the servers doing the work. An internal line of business application serving four hundred staff will barely register.
Inside the cloud there are three more charges people miss. Traffic between availability zones is billed in both directions on the major providers, which turns a chatty microservice architecture spread across zones for resilience into a recurring bill. NAT gateways charge an hourly rate plus a per gigabyte processing fee, so routing all your private subnet traffic to package repositories and object storage through one is an expensive habit; gateway endpoints for object storage remove that path entirely and cost nothing on AWS. And cross region replication is charged as transfer, which is fine when it was a deliberate resilience decision and painful when it was a default.
The other side of this is exit. Egress pricing has historically made leaving a provider expensive, and regulatory pressure in Europe has pushed the major providers to waive transfer charges for customers moving their data off the platform entirely, subject to their own conditions and process. If a future exit matters to your board, check the current terms with the provider and with your counsel rather than relying on a summary in a slide deck, ours included, because these terms have changed more than once.
The other four causes
Sizing carried over from on premise. A physical server bought for a five year peak, running at nine percent CPU, becomes a cloud instance of equivalent specification billed by the hour. On premise that waste was a sunk cost. In the cloud it is a monthly invoice. Right sizing from observed utilisation is the single largest saving available in most migrations and it belongs in the plan before the move, not after the bill.
Non production running all the time. Development, test and staging environments do not need to be up at 03:00 on a Sunday. Scheduled stop and start on non production commonly removes around two thirds of their running hours, and it is a day of work to implement.
Storage sprawl. Snapshots taken during migration and never cleaned up, volumes left attached to deleted instances, log data on the most expensive storage class with no lifecycle rule, and the block storage default that has been superseded by a cheaper and faster generation.
Licensing. Bringing your own licence, licence mobility, and hybrid benefit programmes can change the cost of a Windows or database workload substantially in either direction, and some vendor terms treat cloud hosting differently from on premise in ways that are contentious. Get your licensing position reviewed by someone who reads the actual agreement before you size the estate around an assumption.
Post-Migration Cost Optimisation
The migration ends and the optimisation starts. Doing it in that order is deliberate: right sizing against production traffic you can actually observe beats guessing during a move, and commitment purchases made before the estate has settled lock you into shapes you are about to change.
Measure before you cut
Get the cost and usage data into one place with tags that mean something, then attribute spend to teams and applications. AWS Cost Explorer and the Cost and Usage Report, Azure Cost Management, and Google's billing export to BigQuery all do this; the FinOps Foundation's FOCUS specification is worth adopting if you are multi cloud and tired of reconciling three different schemas. For Kubernetes estates, OpenCost or Kubecost attribute cluster cost down to the namespace, which is the only way to have a useful conversation with a team about their share of a shared cluster.
The order we work in
Right sizing first, using Compute Optimizer, Azure Advisor or Google's recommenders as a starting point and human judgement as the filter, because those recommenders see CPU and memory but not your quarter end. Then scheduling on non production. Then storage lifecycle: intelligent tiering or lifecycle rules on object storage, moving to the current generation of block storage, and deleting the orphaned snapshots.
Only then commitments. Savings Plans, Reserved Instances and committed use discounts return meaningful savings for a one or three year commitment, and they are the right move once the estate is stable. Buying them in month one against pre optimisation sizing is how organisations end up paying for capacity in a shape they no longer run.
After that come the architectural changes: autoscaling that actually scales down, ARM based instance families such as Graviton for workloads whose stack supports them, spot or preemptible capacity for anything interruptible, and serverless for genuinely spiky or low duty cycle services.
Make it a habit, not a project
Cost optimisation done once decays. A monthly review with the cost data in front of the engineering leads, an anomaly alert on unexpected spend, cost visibility in the same dashboard as reliability metrics, and a named owner produces steady results where a heroic quarterly clean up does not. This is the part clients most often ask us to run as an ongoing retainer after the migration itself finishes, and it is a reasonable thing to hand to a team that already knows the estate.
Three Migrations and What Went Wrong in Each
These are composite scenarios drawn from patterns that recur across this kind of work, not accounts of named clients. The failures in them are common enough that if you are mid migration, at least one will be familiar.
The data centre exit with a rewrite attached
A logistics business has a lease ending in fourteen months and around two hundred virtual machines. The architecture team, correctly seeing that the core order system is a monolith that has needed splitting for years, proposes decomposing it during the migration. Nine months in, the data centre exit has a fixed date and the decomposition is at forty percent, with the remaining sixty percent depending on a data model change nobody has finished designing.
What fixes it is separating the deadlines. Rehost the monolith as it is, take the exit date off the critical path, and run the decomposition as its own programme afterwards against a date you control. The cost is a rehosted monolith running for a year longer than the architecture team wanted. The alternative is missing a contractual date and paying for a lease extension, which costs more and is more visible. If your migration currently has a rewrite inside it and an external deadline outside it, this is the conversation to have this month.
The Oracle database that was scoped as a data move
A financial services firm plans to move an Oracle database to PostgreSQL. The project is scoped from the size of the data, around two terabytes, and estimated at three months. Schema conversion tooling converts most of the objects and reports the rest. That report is the actual project: forty thousand lines of PL/SQL including packages with autonomous transactions, hierarchical queries used throughout reporting, and application code that relied on Oracle's default null sorting in ways nobody had documented because it had never mattered.
What fixes it is re-scoping honestly. The data movement stays a three week task. The application change becomes the programme, with an assessment phase that reads the actual PL/SQL and produces a converted, tested equivalent module by module, plus a compatibility layer such as Babelfish where the target and the source language allow it. The reason to raise this before contracting is that a heterogeneous database migration presented as a data project will fail its estimate by a factor, and everyone involved will have known by month two.
The bill that arrived after a clean migration
A software company completes a technically flawless rehost of one hundred and twenty servers. No data loss, no missed windows. The first full month's invoice is well above the on premise run rate. Investigation finds all four of the usual causes at once: instances sized from the old hardware specification, forty non production machines running continuously, several terabytes of migration snapshots never deleted, and an application spread across three availability zones for resilience whose components exchange enough traffic between zones to constitute a visible line item.
What fixes it is an eight week optimisation pass in the order described above, and then a monthly review to stop it recurring. What would have prevented it is target sizing based on observed utilisation during discovery, a snapshot lifecycle policy written into the runbook, non production schedules built during the landing zone phase, and a data transfer model in the original business case. Every one of those is cheap in advance and awkward afterwards, which is exactly why they get skipped.
How We Run a Migration From India
You are hiring a team you will probably never meet in a country you may never visit, to move systems your business depends on. Here is how that actually works, including the parts that are inconvenient.
The overlap window, honestly
India is UTC plus five and a half hours. On a standard 09:30 to 18:30 IST day, the UK gets roughly five hours of live overlap, Australia's eastern states about three, New Zealand about one, and the US East Coast close to zero, because the Indian working day finishes as the American one begins. That is the honest arithmetic and no amount of positive framing changes it.
There are two ways to deal with it. The first is a shifted schedule: a team working 13:30 to 22:30 IST gives US Eastern about four hours of genuine overlap, and a later shift does the same for the US West Coast. That is not free. It costs the engineers their evenings, it narrows the hiring pool because not everyone will take it, and it needs to be agreed at the start rather than requested in month two. The second is to run written first, with a deliberately small live meeting surface, and to accept that most communication is asynchronous by design.
In practice we do both, and we agree the specific window with you before anyone is assigned. Where a shift is required we tell you that up front, because a team that agreed to it reluctantly will not sustain it.
Where the timezone gap helps
Migration is one of the few types of work where the gap is an asset. Your low traffic cutover windows are our working hours. A window starting 22:00 on a Saturday in New York begins at 07:30 on Sunday morning in India, which means the people executing a four hour runbook are at the start of their day, not the fourteenth hour of it. Long running data transfers, replication catch up and validation sweeps also fit naturally into the hours when your business is asleep and nobody is asking for a status update.
The working rhythm
A written daily update posted before your day starts: what moved, what is blocked, what needs a decision from you and by when. One live call in the overlap window, kept short. A weekly session with the migration dashboard open: wave status, servers moved against plan, open risks, and current spend against the model. Everything of consequence written down, because a decision taken verbally at the end of a call is a decision that will be remembered differently by two people in a fortnight.
Infrastructure changes go through pull request review the same way application code does. Two people on every runbook, and the person who writes the runbook is not the person who wrote the automation it describes, which is the fastest way to find the missing step. Nothing goes to production from a console session.
Who you talk to
You talk to the engineers doing the work, not to an account manager relaying messages. English fluency is assessed in a working session during hiring rather than from a certificate: we ask a candidate to explain a technical trade off to a non specialist and to disagree with an interviewer's proposal, because the failure mode that costs you money is an engineer who agrees with a bad instruction rather than one whose accent takes a week to tune into.
Access, security and where your data sits
Engineers work in your cloud accounts under named identities in your directory, with scoped roles and hardware backed multi factor authentication, so every action is attributable in your own audit log. Production data is not copied to local machines. Where the data is sensitive we work through a bastion or a browser based workspace with no local storage, on managed devices with disk encryption and screen lock policy.
Your data stays in the regions you choose. Migration engineers sitting in India does not change where your workloads run or which jurisdiction they sit in, and that distinction matters for a GDPR, HIPAA or state privacy position. Contract scope, IP assignment, confidentiality and data processing terms are set out in the agreement we sign with you before work starts. Where you need alignment to a specific control framework, we work to your control set and evidence it rather than claiming a certification we do not hold. India's Digital Personal Data Protection Act, 2023 applies to us as a processor, which is a floor rather than a substitute for your own obligations, and your counsel should confirm what your regulator expects.
The talent side, plainly
India has a deep pool of engineers with real cloud platform experience, concentrated in Bengaluru, Pune, Hyderabad, Chennai and Mumbai, and a large share of the global delivery centres of the major consultancies sit here, which is where a lot of that experience was built. What is scarcer than the certifications suggest is migration experience specifically: people who have sat through a failed cutover, executed a rollback, and reconciled diverged data afterwards. That is what we screen for, and it is why assembling a migration pod takes longer than assembling a general infrastructure team.
How Do You Keep Quality Under Control From Nine Thousand Kilometres Away?
By making the work visible rather than by trusting reports about it. Four mechanisms, all of which you can inspect yourself at any time.
Everything is in your repositories
Landing zone, network, policy, pipelines and migration automation all live as code in your version control, under review, from day one. Not in an internal repository that gets handed over at the end. You can read every commit as it lands, and if the engagement ends tomorrow you already have everything.
Definition of done, agreed before wave one
A migrated workload is not done because it booted. It is done when it is monitored with alerts routed somewhere a human sees them, backed up with a restore that has been tested, documented in a runbook someone outside the migration team can follow, tagged for cost attribution, reachable only through the intended network path, and signed off by the application owner after their own validation. That checklist is agreed with you before the first wave and applied to every wave after it.
Wave retrospectives that change the runbook
After each wave, what went wrong goes into the runbook as a change, not into a lessons learned document nobody reopens. The measure of whether this is working is that wave four takes fewer hours than wave one for a comparable set of servers. If it does not, something is wrong with the process and that is worth escalating rather than absorbing.
Knowledge does not sit in one head
Two engineers touch every workload, ownership rotates across waves, and runbooks are written by whoever did not build the thing they describe. It is slower in the first month and it is the difference between a team and a single point of failure by the sixth. Attrition in the Indian technology market is real, and pretending otherwise is how clients get surprised; the defence is that the work is in code and the documentation is written by the second pair of hands, so a replacement is a slowdown rather than a restart.
Risks We Name Before You Sign
Access provisioning is the usual first month delay
Not technology. Waiting for VPN accounts, identity federation, security review of offshore access and a decision on whether an external engineer may see production data. Started after signature, it can absorb three weeks of a paid team's time. We begin it during contracting, and where production access genuinely cannot be granted early we work against a discovery dataset and the landing zone build rather than sitting idle.
Discovery finds things that change the plan
It usually does. An undocumented dependency on a partner's network, a licence tied to hardware, an application whose vendor no longer exists. The plan produced before discovery is an estimate, and we say so, with a revision point built in after the dependency map is complete rather than defending a number we produced with less information.
The hybrid period is longer than planned
Phased migrations run with a foot in both worlds, and that period tends to extend because the last few applications are always the hardest. The network link, the duplicated monitoring and the dual operational burden all have a running cost. We put a target end date on the hybrid period and report against it, because an indefinite hybrid state is the most expensive place an estate can sit.
Your own team's time is a real cost
Application owners have to validate their systems after each wave, and that cannot be outsourced to us because we do not know what correct looks like in your business. Expect a meaningful commitment from a senior person in weeks one and two, then application owners for a few hours per wave, plus someone empowered to make decisions inside the overlap window. Migrations stall on decision latency more often than on engineering.
Ramp up is not day one
A new team on an unfamiliar estate reaches full productivity in roughly four to six weeks, and a complex estate takes longer. That ramp is a genuine cost and it should be in your plan. Anyone promising full velocity from week one has not told you about it, which does not mean it is not there.
If it does not work out
Exit arrangements belong in the agreement from the start rather than being negotiated when relations are strained. Notice, handover length and the knowledge transfer process are all set out in the contract before work begins, and everything sits in your repositories and your cloud accounts throughout, so there is nothing physically to hand back. Offboarding runs to a documented access revocation checklist. You should never be in a position where leaving us is technically difficult; if you were, we built it wrong.
Engagement Models
Dedicated migration pod
A named team, typically a lead cloud engineer, two migration engineers and a database specialist, with a fractional architect, working only on your programme. Right when the estate is large enough for waves to run continuously for months. The same people from wave one to decommissioning, not a rotating bench, on terms agreed before the pod starts.
Scoped assessment and plan
Discovery, dependency mapping, per application R decisions, wave plan, target sizing and a cost model, delivered as documents your own team can execute against or take to another supplier. Right when you need a defensible plan and a number before committing to a programme. Frequently the first thing we do, and it stands alone.
Wave delivery or post-migration retainer
Either execution of an agreed set of waves against a plan that already exists, or ongoing operation after the move: cost optimisation, patching, monitoring and incident response within agreed hours. Two situations lead here. You have the plan but not the hands, or the migration is finished and you would rather not hire permanently for the operate phase.
Where This Sits Alongside Our Other Work
Migration is one phase of a longer relationship with a cloud platform. The target architecture, the pipelines, the container platform and the reliability practice around it are separate disciplines, and we keep them separate deliberately so a migration deadline does not swallow them. Our broader cloud services practice covers the design and operation of the platform your workloads land in, and DevOps engineering covers the pipelines, infrastructure as code and container work that a migrated estate needs next.
Where the reason for moving is an application nobody wants to touch, legacy application modernization is the honest framing rather than a migration, and it should be scoped as such. Where the data platform is the thing being moved, data engineering covers the pipelines and warehouse side. For staffing rather than a scoped programme, you can hire AWS developers in India for platform work, or hire DevOps developers in India to strengthen a team you already have.
Frequently Asked Questions About Cloud Migration in India
How long does a cloud migration take?
Discovery and dependency mapping is four to six weeks for a mid sized estate, and it is the phase that produces an honest estimate for everything after it. A straight rehost of fifty servers with no database version changes typically runs three to five months from landing zone to decommissioning. Add a heterogeneous database migration, such as Oracle to PostgreSQL, and you are into nine to eighteen months because the application code has to change with it. Anyone quoting the whole programme before seeing your dependency map is guessing.
Why is the first cloud bill so much higher than the estimate?
Usually four reasons stacked together. Servers were sized on their on premise specification rather than their actual utilisation, so a machine running at eight percent CPU became an equally large instance. Data transfer was never modelled, so cross zone chat, NAT gateway processing and egress to the internet appear as line items nobody forecast. Non production environments run all night and all weekend because nobody built a schedule. And snapshots accumulate with no lifecycle rule. All four are fixable, and all four are cheaper to design out than to discover.
Is lift and shift a mistake?
No, but treating it as the finish line is. Rehosting is the right call when you have a data centre exit date, when the application is stable and unchanging, or when you need to stop a hardware refresh. It is the wrong call for anything you were about to rewrite anyway, for licence bound workloads where the cloud licensing terms are worse, and for applications whose cost problem is architectural. The failure mode is not the rehost itself. It is stopping there, keeping the same over provisioned shapes, and concluding that the cloud is expensive.
How do you migrate a database with no downtime?
You do not get zero downtime. You get a cutover measured in seconds or minutes instead of hours. A tool such as AWS DMS, Azure Database Migration Service or native logical replication takes a full load, then applies change data capture until the replica is within a second of the source. At the cutover you stop writes, wait for lag to reach zero, verify row counts and checksums, repoint the application, and start writes again. The prerequisites are the part people miss: supplemental logging, primary keys on every replicated table, and a plan for the object types the tool silently does not carry across.
What happens if the cutover goes wrong at three in the morning?
You roll back, which only works if it was designed before the cutover rather than improvised during it. That means reverse replication configured and tested from the new database back to the old one, the source environment kept running rather than decommissioned, DNS time to live lowered days in advance so a repoint actually propagates, and a written point of no return. After the first write lands on the new system and does not replicate backwards, rollback stops being a switch and becomes a data reconciliation exercise. We rehearse the rollback in a dry run, not just the migration.
Can an offshore team run the cutover night for a US business?
For US clients this is one of the few places the timezone gap works in your favour. A low traffic cutover window of 22:00 Saturday US Eastern is 07:30 Sunday in India, so the engineers running the migration are fresh at the start of their day rather than fourteen hours into it. What has to be agreed in advance is who on your side is awake and authorised to make the abort call, because that decision cannot sit with the team executing the runbook.
What is the real overlap window with an engineering team in India?
On a standard 09:30 to 18:30 IST day, the UK gets about five hours of overlap, Australian eastern states about three, New Zealand roughly one, and US Eastern effectively none, because the Indian working day ends as yours begins. A shifted 13:30 to 22:30 IST schedule gives US Eastern about four hours of genuine overlap. That shift is a real cost in coordination and in people's evenings, so we agree the window with you up front rather than claiming round the clock coverage comes free.
Do we need a landing zone before we migrate anything?
You need the parts that are expensive to change later: the account or subscription structure, network address ranges, identity federation, logging destinations and a tagging standard. Those are load bearing, and retrofitting them across forty accounts after the fact is a project of its own. You do not need every guardrail and every policy in place before the first workload moves. Build the skeleton with AWS Control Tower, Azure Landing Zones or the equivalent, migrate a low risk application through it, then harden with what you learned.