smoke test: support ceph storage pools - #11931
weizhouapache wants to merge 6 commits into
Conversation
|
@blueorangutan package |
|
@weizhouapache a [SL] Jenkins job has been kicked to build packages. It will be bundled with KVM, XenServer and VMware SystemVM templates. I'll keep you posted as I make progress. |
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #11931 +/- ##
============================================
+ Coverage 19.78% 21.18% +1.40%
- Complexity 19995 20200 +205
============================================
Files 6371 5886 -485
Lines 575909 535238 -40671
Branches 70509 62754 -7755
============================================
- Hits 113950 113397 -553
+ Misses 449526 409503 -40023
+ Partials 12433 12338 -95
Flags with carried forward coverage won't be shown. Click here to find out more. ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
|
Packaging result [SF]: ✔️ el8 ✔️ el9 ✔️ el10 ✔️ debian ✔️ suse15. SL-JID 15584 |
|
@blueorangutan test matrix |
|
@weizhouapache a [SL] Trillian-Jenkins matrix job (EL8 mgmt + EL8 KVM, Ubuntu22 mgmt + Ubuntu22 KVM, EL8 mgmt + VMware 7.0u3, EL9 mgmt + XCP-ng 8.2 ) has been kicked to run smoke tests |
|
[SF] Trillian test result (tid-14741)
|
|
[SF] Trillian test result (tid-14742)
|
|
@blueorangutan package |
|
@weizhouapache a [SL] Jenkins job has been kicked to build packages. It will be bundled with KVM, XenServer and VMware SystemVM templates. I'll keep you posted as I make progress. |
|
Packaging result [SF]: ✔️ el8 ✔️ el9 ✔️ el10 ✔️ debian ✔️ suse15. SL-JID 15597 |
|
@blueorangutan test |
|
@weizhouapache a [SL] Trillian-Jenkins test job (ol8 mgmt + kvm-ol8) has been kicked to run smoke tests |
|
[SF] Trillian test result (tid-14744)
|
|
[SF] Trillian test result (tid-14753)
|
|
[SF] Trillian test result (tid-14743)
|
ddfdacf to
844de76
Compare
|
ready for review and testing |
|
@blueorangutan package |
|
@blueorangutan package |
|
@sureshanaparti a [SL] Jenkins job has been kicked to build packages. It will be bundled with KVM, XenServer and VMware SystemVM templates. I'll keep you posted as I make progress. |
|
Packaging result [SF]: ✔️ el8 ✔️ el9 ✔️ el10 ✔️ debian ✔️ suse15. SL-JID 16825 |
test_backup_recovery_nas.py only allowed NFS primary storage, since it reused the primary storage pool's own path as the NAS backup repository address, and always required incremental-backup semantics that only qcow2/NFS storage can provide. Neither holds on Ceph/RBD. - setUpClass now accepts RBD alongside NFS as the primary storage pool type, picking a pool that's actually Up rather than list()[0] -- environments that added Ceph/RBD after the zone's original NFS primary storage keep that old pool around in Disabled state, and it still sorts first, silently exercising its path as if it were the storage VMs actually deploy on. - The NAS backup repository's NFS export address is resolved independently of the primary storage: when the primary pool isn't NFS, reuse the nfs test data entry (services[nfs][url]) -- the same temporary NFS mount point test_primary_storage.py uses for its temporary NFS primary storage pool, and something every marvin environment already has configured. An explicit nas_backup_repository_address test data entry or NAS_BACKUP_REPO_ADDRESS environment variable, if set, takes precedence. - The external offering imported in setUpClass is matched to the repository just created by externalid (== the repository's own id for the nas provider) rather than blindly taking index 0 -- a stray repository left over from an earlier interrupted run, whose backups didn't get cleaned up so its own teardown couldn't remove it either, sorts alongside the new one with no guarantee of which comes first. - Incremental NAS backups require QEMU dirty bitmaps / libvirt checkpoints, which only exist on file-based qcow2 storage (NASBackupProvider.allVolumesOnCheckpointCapableStorage). The six incremental-chain tests now skip on RBD/Ceph, where the provider always falls back to full-only backups server-side, rather than failing on assertions that storage type can never satisfy. - Added test_restore_volume_and_attach_to_vm, which exercises restoreVolumeFromBackupAndAttachToVM end-to-end (restoring a backed-up ROOT and DATADISK volume onto a second, stopped Instance) -- the API that drives the restore-and-attach code fixed by the previous commit (apache#14007). The target Instance is stopped with forced=True: a graceful ACPI stop was observed to time out (~2 minutes) before falling back to a hard destroy anyway, and once forced to a hard destroy the domain drops out of libvirt entirely, so the periodic ping-based PowerState sync the restore call depends on falls back to a much slower heuristic well past any reasonable wait. A forced stop destroys the domain immediately and deterministically.
844de76 to
fff8e56
Compare
A zone-wide storage pool (e.g. RBD configured with scope=ZONE) isn't bound to a single pod/cluster, so a volume on it has no clustername/ clusterid/podid/podname. test_10_list_volumes asserted these were always non-None, which fails whenever storage is zone-scoped rather than cluster-scoped. Now look up the volume's actual storage pool and only assert pod/cluster attributes when its scope isn't ZONE.
When a host has no explicit tags initially, cls.host.hosttags is None. Passing None to the API does not set the hosttags parameter, so tags are never cleared during cleanup. Convert None to empty string to properly clear tags that were set during the test.
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Storage selection can incorrectly run or skip tests, and failed restore tests can leave backups behind.
Review effort: Balanced
Findings: 6
Open (6)
Restrict storage pool selection to eligible pools in the test zone · New Use actual attached volume pools for backup capability checks · New Delete created backups during failure cleanup · New Require an active pool before selecting the test candidate · New Gate snapshot tests using the deployed VM's volume pool · New Check encryption support in the selected test zone · New
What changed in this PR
Extends CloudStack’s smoke tests to support Ceph/RBD storage while skipping unsupported operations.
Changes:
- Enables RBD coverage for direct downloads, overprovisioning, snapshots, and NAS backups.
- Adds storage-aware checks and skips.
- Adds offline backup-volume restore-and-attach coverage.
| File | Description |
|---|---|
| test/integration/smoke/test_volumes.py | Adjusts zone-wide pool assertions and encryption skips. |
| test/integration/smoke/test_vm_snapshots.py | Adds an RBD snapshot skip. |
| test/integration/smoke/test_vm_life_cycle.py | Skips live migration between RBD pools. |
| test/integration/smoke/test_snapshots.py | Allows RBD snapshot coverage. |
| test/integration/smoke/test_over_provisioning.py | Includes RBD overprovisioning tests. |
| test/integration/smoke/test_host_tags.py | Sends empty strings to clear host tags. |
| test/integration/smoke/test_direct_download.py | Includes RBD in shared-storage coverage. |
| test/integration/smoke/test_backup_recovery_nas.py | Adds Ceph support, repository selection, incremental skips, and restore coverage. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
| storage_pools = StoragePool.list(cls.api_client) | ||
| usable_pools = [p for p in storage_pools if getattr(p, 'state', 'Up') == 'Up'] | ||
| cls.storage_pool = usable_pools[0] if usable_pools else storage_pools[0] | ||
| if cls.storage_pool.type.lower() not in SUPPORTED_PRIMARY_STORAGE_POOL_TYPES: | ||
| cls.skipTest(cls, reason="Test can be run only if the primary storage is of type NFS or RBD (Ceph)") |
| if self.storage_pool.type.lower() == 'rbd': | ||
| self.skipTest("Incremental backups are not supported on RBD/Ceph primary Storage") |
| Backup.delete(self.apiclient, backup.id) | ||
| finally: |
| ) | ||
| for pool in storage_pools: | ||
| if not cls.nfsStorageFound and pool.type == "NetworkFilesystem": | ||
| if not cls.nfsStorageFound and pool.type in ("NetworkFilesystem", "RBD"): |
| list_volume_pool_response = list_storage_pools(cls.apiclient) | ||
| volume_pool = list_volume_pool_response[0] | ||
| if volume_pool.type == "RBD": | ||
| cls.skipTest(cls, reason="VM snapshot is unsupported for VMs on RBD storage pool") |
| list_volume_pool_response = list_storage_pools(cls.apiclient) | ||
| volume_pool = list_volume_pool_response[0] | ||
| if volume_pool.type == "RBD": | ||
| cls.skipTest(cls, reason="Volume encryption is unsupported for volumes on RBD storage pool") |

Description
This PR improves and fixes some tests on ceph storage
Types of changes
Feature/Enhancement Scale or Bug Severity
Feature/Enhancement Scale
Bug Severity
Screenshots (if appropriate):
How Has This Been Tested?
How did you try to break this feature and the system with this change?