Skip to content

DAOS-19518 test: nvme/fault.py - Update timeout and delay - #18898

Draft
shimizukko wants to merge 1 commit into
masterfrom
makito/DAOS-19518
Draft

DAOS-19518 test: nvme/fault.py - Update timeout and delay#18898
shimizukko wants to merge 1 commit into
masterfrom
makito/DAOS-19518

Conversation

@shimizukko

Copy link
Copy Markdown
Contributor

test_nvme_fault sets NVMe devices faulty. If the device is an sysXs device, "storage set nvme-faulty" command would fail, which is expected.

The issue is that if we call dmg pool query (to check rebuild status) immeidately after the nvme-faulty command failure, the pool query command may timeout (after 5 min.). If pool query times out, we should retry. To enable the retry, set pool_query_timeout larger than 5 min, say 330 sec.

Another way that could avoid this issue is to add some sleep between nvme-faulty and pool query. We already have a variable for this in TestPool, pool_query_delay, so set some time to it in the test yaml and move the sleep code in query() to right before calling pool query.

Skip-fault-injection-test: true
Skip-func-hw-test-large: false
Test-tag: test_nvme_fault
Test-repeat: 3

Steps for the author:

  • Commit message follows the guidelines.
  • Appropriate Features or Test-tag pragmas were used.
  • Appropriate Functional Test Stages were run.
  • At least two positive code reviews including at least one code owner from each category referenced in the PR.
  • Testing is complete. If necessary, forced-landing label added and a reason added in a comment.

After all prior steps are complete:

  • Gatekeeper requested (daos-gatekeeper added as a reviewer).

test_nvme_fault sets NVMe devices faulty. If the device
is an sysXs device, "storage set nvme-faulty" command would
fail, which is expected.

The issue is that if we call dmg pool query (to check
rebuild status) immeidately after the nvme-faulty command
failure, the pool query command may timeout (after 5 min.).
If pool query times out, we should retry. To enable the
retry, set pool_query_timeout larger than 5 min, say 330
sec.

Another way that could avoid this issue is to add some
sleep between nvme-faulty and pool query. We already have
a variable for this in TestPool, pool_query_delay, so
set some time to it in the test yaml and move the sleep
code in query() to right before calling pool query.

Skip-fault-injection-test: true
Skip-func-hw-test-large: false
Test-tag: test_nvme_fault
Test-repeat: 3
Signed-off-by: Makito Kano <makito.kano@hpe.com>
@github-actions

Copy link
Copy Markdown

Ticket title is 'nvme/fault.py:NvmeFault.test_nvme_fault - pool query timeout detecting rebuild'
Status is 'Open'
Labels: 'ci_master_weekly,testp2,weekly_test'
https://daosio.atlassian.net/browse/DAOS-19518

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Development

Successfully merging this pull request may close these issues.

1 participant