All notable changes to this project will be documented in this file.
The format is based on Keep a Changelog, and this project adheres to Semantic Versioning.
Release: https://crates.io/crates/duroxide/0.1.20
- Custom status as history events —
set_custom_status()andreset_custom_status()now emitCustomStatusUpdatedhistory events instead of writing toExecutionMetadata.custom_status. Custom status is now fully durable, replayable, and deterministic across turns.- Removed
CustomStatusUpdateenum from providers - Removed
ExecutionMetadata.custom_statusfield - Provider
ack_orchestration_item()must now scanhistory_deltafor the lastCustomStatusUpdatedevent and apply it to the instances table (see provider-implementation-guide.md) initial_custom_statusfield added toOrchestrationStartedevent andContinueAsNewwork item for carry-forward across continue-as-new boundaries
- Removed
get_custom_status()on OrchestrationContext — Read the current custom status value, reflecting allset_custom_status/reset_custom_statuscalls across turns and CAN boundariesshort_poll_threshold()on ProviderFactory — Configurable timing for short polling validation tests; remote-database providers can override with higher values (closes #51)test_orphan_activity_after_instance_force_deletion— Provider validation test verifying graceful handling of activities orphaned by instance force-deletion (closes #37)
test_cancelling_nonexistent_activities_is_idempotentnow usesexecution_id: 1instead of99to correctly validate same-execution cancellation semantics (closes #40)- Removed dead
ActivityContext::newconstructor (unused, runtime usesnew_with_cancellation) - Removed unused
clippy::clone_on_ref_ptrsuppression from observability.rs (closes #48)
Release: https://crates.io/crates/duroxide/0.1.19
Proposal: Custom Status Progress Proposal: External Event Semantics Proposal: Persistent Event Queuing
-
Event Queue API — Persistent FIFO event queues that survive
continue_as_newctx.dequeue_event(queue)andctx.dequeue_event_typed::<T>(queue)for orchestrationsclient.enqueue_event(instance, queue, data)andclient.enqueue_event_typed::<T>()for clients- FIFO ordering with buffering — messages can arrive before orchestration subscribes
- Queue messages carry forward across continue-as-new boundaries
-
Custom Status — Orchestration progress reporting visible to external clients
ctx.set_custom_status(json)publishes structured progress from orchestrationsclient.wait_for_status_change(instance, version, poll, timeout)for efficient polling- Status persists across continue-as-new boundaries
- New Provider trait methods:
set_custom_status(),get_custom_status() - Schema migration
20240108000000_add_custom_status.sql
-
Retry on Session — Combine retry policies with session affinity (closes #56)
ctx.schedule_activity_with_retry_on_session(name, input, policy, session_id)ctx.schedule_activity_with_retry_on_session_typed::<In, Out>()- All retry attempts pinned to the same worker session
-
Typed event helpers —
client.raise_event_typed::<T>()andclient.enqueue_event_typed::<T>() -
Provider validation —
test_prune_bulk_includes_running_instancescatches providers that exclude Running instances from bulk prune (closes #50) -
Scenario test — Copilot Chat pattern: multi-turn chat using dequeue_event + set_custom_status + CAN
- Renamed persistent event internals:
ExternalRaisedPersistent→QueueMessage,ExternalSubscribedPersistent→QueueSubscribed
client.raise_event_persistent()— useclient.enqueue_event()insteadctx.schedule_wait_persistent()— usectx.dequeue_event()instead
Release: https://crates.io/crates/duroxide/0.1.18
Proposal: Activity Implicit Sessions v2
-
Activity Session Affinity — Route activities to the same worker for in-memory state reuse
ctx.schedule_activity_on_session(name, input, session_id)pins activities by session IDctx.schedule_activity_on_session_typed()for serde-based typed inputs/outputsActivityContext::session_id()getter for process-local state lookup- Two-timeout model:
session_lock_timeout(heartbeat lease) +session_idle_timeout(inactivity expiry) - Automatic session lifecycle: implicit creation, heartbeat renewal, idle unpin, crash recovery
SessionTrackerenforcesmax_sessions_per_runtimeacross all worker slots via RAII guardsworker_node_idoption for stable session identity across restarts (e.g., K8s StatefulSet pods)- Session manager background task for lock renewal and orphan cleanup
-
Provider API changes (required)
fetch_work_item()gainssession: Option<&SessionFetchConfig>parameter for session routing- New
renew_session_lock()method — batched heartbeat for owned non-idle sessions - New
cleanup_orphaned_sessions()method — sweep expired session rows with no pending work ack_work_item()andrenew_work_item_lock()piggybacklast_activity_atupdates (guarded bylocked_until)
-
New
RuntimeOptionsfieldssession_lock_timeout(default 30s),session_lock_renewal_buffer(default 5s)session_idle_timeout(default 5min),session_cleanup_interval(default 5min)max_sessions_per_runtime(default 10),worker_node_id(default None)
-
33 provider validation tests for session routing, locks, races, and cross-concern interactions
-
22 E2E tests for single/multi-worker, fan-out, CAN, heterogeneous scenarios
-
Schema migration
20240107000000_add_sessions.sql— sessions table + worker_queue.session_id
session_id: Option<String>added toAction::CallActivity,EventKind::ActivityScheduled, andWorkItem::ActivityExecute(backward compatible viaserde(default))- Existing
schedule_activitycalls are completely unaffected (session_id = None) - Documentation updated across ORCHESTRATION-GUIDE, provider-implementation-guide, provider-testing-guide, README, and migration-guide
Release: https://crates.io/crates/duroxide/0.1.17
Proposal: Provider Capability Filtering
-
Provider Capability Filtering (Phase 1) — Safe rolling upgrades in mixed-version clusters
- Orchestration dispatcher passes a version filter to the provider so it only returns
executions whose pinned
duroxide_versionfalls within the runtime's supported range - SQL-level filtering applied before lock acquisition and history deserialization
- NULL pinned version treated as always compatible (backward compat with pre-migration data)
RuntimeOptions::supported_replay_versionsfor custom version range configuration- Defense-in-depth: runtime-side compatibility check after fetch with 1-second abandon delay
- Startup log declaring supported version range; warning log on incompatible-version abandon
- Orchestration dispatcher passes a version filter to the provider so it only returns
executions whose pinned
-
New types:
SemverVersion,SemverRange,DispatcherCapabilityFilter,current_build_version() -
Provider API change:
fetch_orchestration_item()gainsfilter: Option<&DispatcherCapabilityFilter>parameter -
History deserialization contract — Providers must surface deserialization errors (not silently drop events)
history_errorfield on fetched items for deserialization failures- Transaction commits lock + attempt_count before returning errors (enables poison path)
-
ProviderFactory test helpers —
corrupt_instance_history()andget_max_attempt_count()optional methods for provider-agnostic deserialization contract tests -
38 new tests — 20 provider validation + 18 e2e scenario tests covering filtering, rolling deployment routing, metadata/migration, ContinueAsNew isolation, drain procedures, and observability
- Migration:
20240106000000_add_pinned_version.sqladdsduroxide_version_major/minor/patchcolumns toexecutionstable - Provider validation test total: 114 tests (up from 94)
Release: https://crates.io/crates/duroxide/0.1.16
Proposal: Persist Cancellation Decisions in History
-
Cancellation History Events - Record cancellation decisions as durable history breadcrumbs
- New
ActivityCancelRequestedandSubOrchestrationCancelRequestedevent kinds - Dropped futures (select losers, terminal cleanup) now recorded in history
- Enables observability: history answers "was this cancelled?"
- Enables replay determinism: detect when cancellation decisions differ on replay
- Idempotent: side-channel cancellations only emitted once per decision
- New
-
Nondeterminism Tests for Completion Validation - 12 new tests covering:
- Completion kind mismatches (timer/activity/sub-orchestration cross-checks)
- Duplicate completion detection for closed schedules
- OrchestrationChained mismatch validation
- Duplicate Completion Detection - Fixed oversight where
open_schedules.remove()was never called after delivering completions in replay engine- Previously, duplicate completions in history were silently accepted
- Now properly triggers nondeterminism error on duplicate completions
- Housekeeping - Moved implemented/rejected proposals to
docs/proposals-impl/:metrics-facade-migration.md(implemented)replay-simplification-PROGRESS.md(completed)activity-cancellation-queue-flag.md(rejected/superseded by lock-stealing)
Release: https://crates.io/crates/duroxide/0.1.15
- Simplified Metrics Facade - Internal observability uses consistent atomic counters with a cleaner facade pattern
- Code Coverage Improvements - Test coverage improved to 91.9% with better organization
- New provider validation tests for error handling, management interface, and observability
- Removed duplicate tests that overlapped with provider validation suite
- Code Coverage Guide - New
docs/code-coverage-guide.mdwith llvm-cov setup instructions - Copilot Skill for Coverage - AI assistant skill for code coverage workflows
- README Markdown Formatting - Fixed section heading syntax for better rendering
Release: https://crates.io/crates/duroxide/0.1.14
- Fire-and-Forget Orchestrations Now Record History Events -
ctx.schedule_orchestration()(detached/chained orchestrations) now correctly createsOrchestrationChainedevents in history- Previously, these fire-and-forget calls were not recorded, breaking determinism detection on replay
- If an orchestration scheduled a detached orchestration followed by an activity, replay would fail with nondeterminism error
- Added proper action-to-event conversion and event matching in replay engine
- Action-to-Event Recording Tests - New test module
tests/replay_engine/action_to_event.rsverifying all scheduling actions create corresponding history events - E2E Test for Detached + Activity Pattern -
sample_detached_then_activity_fsvalidates the fix end-to-end
Release: https://crates.io/crates/duroxide/0.1.13
Proposal: System Calls as Real Activities
-
System Calls Reimplemented as Regular Activities -
ctx.new_guid()andctx.utc_now()now use normal activity infrastructure- Simplifies replay engine by removing special-case SystemCall handling
- Fixes determinism bugs where syscalls returned fresh values on replay
- Reserved activity prefix
__duroxide_syscall:prevents user collisions - Builtin activities injected automatically at runtime startup
-
API Rename:
utcnow()→utc_now()- Consistent with Rust naming conventions
- Reserved Activity Prefix Validation -
ActivityRegistryrejects names starting with__duroxide_syscall: - Comprehensive Syscall Tests - Replay determinism, ordering, single-thread mode, cancellation
SystemCallvariants fromAction,EventKind,CompletionResult- SystemCall handling from replay engine (no more re-poll loop)
EVENT_TYPE_SYSTEM_CALLfrom sqlite provider
- Updated ORCHESTRATION-GUIDE and durable-futures-internals for new syscall semantics
- Reorganized proposals: moved 11 implemented proposals to
docs/proposals-impl/ - Updated merge prompt to require squash-only merges
utcnow()renamed toutc_now()- update all call sites- Histories containing
SystemCallevents will not replay (pre-1.0, acceptable)
Release: https://crates.io/crates/duroxide/0.1.12
-
Unobserved Future Cancellation - Futures that are scheduled but never awaited are now properly cancelled
- New
DurableFutureimplementation with proper drop semantics - Cancellation events recorded in history for deterministic replay
- Comprehensive test coverage in
tests/replay_engine/unobserved_futures.rs
- New
-
AI Skills System - New
docs/skills/folder for AI coding assistant context- Installation instructions for VS Code Copilot, Claude Code, and Cursor
duroxide-provider-implementationskill for provider developers
-
Provider Validation - New cancellation validation tests
test_activity_cancellation_via_lock_stealing- Additional lock stealing edge case tests
-
Major Documentation Refactor
- Rewrote
provider-implementation-guide.mdwith better structure and pedagogy - Rewrote
architecture.mdwith cleaner ASCII diagrams (removed mermaid) - Merged
replay-engine.mdintodurable-futures-internals.md - Added long-polling vs short-polling explanation
- Added Performance Considerations and ProviderAdmin sections
- Rewrote
-
Simplified ActivityRegistry API - Now takes value instead of Arc
-
Improved Dispatcher Backoff Logic - Better stale activity handling
- Polling Model Documentation - Corrected to multi-poll (not single-poll) model
| Area | Files | Insertions | Deletions | Net |
|---|---|---|---|---|
| Core (src/) | 15 | +2,355 | -1,999 | +356 |
| Tests (tests/) | 63 | +7,616 | -3,353 | +4,263 |
| Docs (docs/) | 43 | +17,477 | -4,004 | +13,473 |
| Total | 134 | +17,459 | -9,536 | +7,923 |
Release: https://crates.io/crates/duroxide/0.1.11
-
Issue #49: WorkItemReader version extraction during completion-only replay
Fixed a bug where
WorkItemReaderdid not extract all fields from history during completion-only replay (when noStartorContinueAsNewitem is present in the work item batch). This caused nondeterminism errors when:- A versioned orchestration was started
- The runtime restarted (or lock expired) mid-execution
- Activity completion arrived without a start item
- The runtime incorrectly used the
Latestversion policy instead of the version recorded in history
The fix extracts all tuple fields (
orchestration_name,input,version,parent_instance,parent_id) fromHistoryManagerduring completion-only replay, ensuring deterministic handler resolution.
- New scenario tests for issue #49 regression prevention:
e2e_replay_completion_only_must_use_version_from_history- First execution replaye2e_replay_completion_only_after_can_must_use_version_from_history- Nth execution (after CAN) replay- Unit tests verifying all
WorkItemReadertuple fields are correctly extracted
Release: https://crates.io/crates/duroxide/0.1.10
-
Rolling Deployment Support - Exponential backoff for unregistered handlers
Unregistered orchestrations and activities now use exponential backoff instead of immediate failure, enabling graceful rolling deployments in multi-node clusters:
- Messages abandoned with backoff (1s → 2s → 4s → ... up to 60s max)
- Bounce between nodes until one with the handler registered picks it up
- Eventually fail as
ErrorDetails::Poisonif handler never becomes available - Configurable via
UnregisteredBackoffConfig(defaults: 1s base, 60s max, 6 exponent cap)
-
New scenario tests for rolling deployments
e2e_rolling_deployment_new_activity- Multi-node deployment with new activitye2e_rolling_deployment_version_upgrade- Version upgrade via continue-as-new
-
Consolidated unregistered handler tests in
tests/unregistered_backoff_tests.rsunknown_version_fails_with_poison- Version mismatch handlingcontinue_as_new_to_missing_version_fails_with_poison- CAN to missing versiondelete_poisoned_orchestration- Cleanup after poison- Plus existing backoff behavior tests
- BREAKING:
ConfigErrorKind::MissingVersionremoved - unregistered handlers now use backoff/poison path config_errormetric now only tracks nondeterminism (unregistered handlers result inpoison)- Updated
docs/metrics-specification.mdwith new error type behaviors - Updated
docs/ORCHESTRATION-GUIDE.mderror handling section
tests/unknown_activity_tests.rs- consolidated intounregistered_backoff_tests.rstests/unknown_orchestration_tests.rs- consolidated intounregistered_backoff_tests.rs
Release: https://crates.io/crates/duroxide/0.1.9
-
Management API for Instance Deletion and Pruning - Comprehensive instance lifecycle management
Client API:
delete_instance(id, force)- Delete single instance with cascadingdelete_instance_bulk(filter)- Bulk delete with filters (IDs, timestamp, limit)prune_executions(id, options)- Prune old executions from long-running instancesprune_executions_bulk(filter, options)- Bulk prune across multiple instancesget_instance_tree(id)- Inspect instance hierarchy before deletion
Provider API (ProviderAdmin trait):
delete_instance(id, force)- Provider-level single deletiondelete_instance_bulk(filter)- Provider-level bulk deletiondelete_instances_atomic(ids)- Atomic batch deletion for cascadingprune_executions(id, options)- Provider-level pruningprune_executions_bulk(filter, options)- Provider-level bulk pruningget_instance_tree(id)- Provider-level tree traversallist_children(id)- List direct child sub-orchestrationsget_parent_id(id)- Get parent instance ID
Safety Guarantees:
- Running instances protected (skip or error based on API)
- Current execution never pruned
- Sub-orchestrations cannot be deleted directly (must delete root)
- Atomic cascading deletes (all-or-nothing)
- Force delete available for stuck instances
-
102 new provider validation tests - Deletion, bulk deletion, pruning, cascading deletes, filter combinations, safety tests
- Provider implementation guide with deletion/pruning contracts
- Provider testing guide updates
- Continue-as-new docs with pruning section
- README instance management section
- Enhanced management-api-deletion proposal with force delete semantics
Release: https://crates.io/crates/duroxide/0.1.8
-
Lock-stealing activity cancellation - New mechanism for cancelling in-flight activities
- Activities are cancelled by deleting their worker queue entries ("lock stealing")
- Workers detect cancellation when lock renewal fails (entry missing)
- More efficient than polling execution state on every renewal
- Enables batch cancellation of multiple activities atomically
-
ScheduledActivityIdentifier- New struct for identifying activities in worker queue- Fields:
instance(String),execution_id(u64),activity_id(u64) - Used by
ack_orchestration_itemto specify activities to cancel - Exported from
duroxide::providers
- Fields:
-
Provider validation tests for lock-stealing - 5 new tests
test_cancelled_activities_deleted_from_worker_queue- Verify deletion during acktest_ack_work_item_fails_when_entry_deleted- Verify permanent error on stolen locktest_renew_fails_when_entry_deleted- Verify renewal fails on stolen locktest_cancelling_nonexistent_activities_is_idempotent- Verify no error for missing entriestest_batch_cancellation_deletes_multiple_activities- Verify batch deletion
-
Worker queue activity identity columns - Store activity identity for cancellation
- New migration:
20240104000000_add_worker_activity_identity.sql - SQLite provider stores
instance_id,execution_id,activity_idon ActivityExecute items
- New migration:
-
BREAKING:
Provider::ack_orchestration_itemsignature changed- Added 7th parameter:
cancelled_activities: Vec<ScheduledActivityIdentifier> - Provider must delete matching worker queue entries atomically in same transaction
- Added 7th parameter:
-
BREAKING:
Provider::fetch_work_itemreturn type simplified- Changed from
(WorkItem, String, u32, ExecutionState)to(WorkItem, String, u32) - Removed
ExecutionState- cancellation detected via lock renewal failure instead
- Changed from
-
BREAKING:
Provider::renew_work_item_lockreturn type changed- Changed from
Result<ExecutionState, ProviderError>toResult<(), ProviderError> - Failure indicates lock was stolen (activity cancelled) or expired
- Changed from
-
BREAKING:
Provider::ack_work_itemmust fail when entry missing- Returns permanent error if work item entry was deleted (lock stolen)
- Signals to worker that activity was cancelled
-
Provider validation test count: 80 tests (up from 75)
ExecutionStateenum removed from Provider API - No longer needed- Was used for state-polling cancellation approach
- Lock-stealing provides more efficient cancellation mechanism
- Provider validation tests for ExecutionState still exist (legacy support during migration)
Provider implementers - Required changes:
- Update
ack_orchestration_itemsignature:
async fn ack_orchestration_item(
&self,
lock_token: &str,
execution_id: u64,
history_delta: Vec<Event>,
worker_items: Vec<WorkItem>,
orchestrator_items: Vec<WorkItem>,
metadata: ExecutionMetadata,
cancelled_activities: Vec<ScheduledActivityIdentifier>, // NEW
) -> Result<(), ProviderError>;- Update
fetch_work_itemreturn type:
async fn fetch_work_item(...) -> Result<Option<(WorkItem, String, u32)>, ProviderError>;
// Removed ExecutionState from tuple- Update
renew_work_item_lockreturn type:
async fn renew_work_item_lock(...) -> Result<(), ProviderError>;
// Returns () instead of ExecutionState- Update
ack_work_itemto fail on missing entry:
// Return error if entry not found (lock was stolen)
if rows_affected == 0 {
return Err(ProviderError::permanent("ack_work_item", "Entry not found (lock stolen)"));
}- Store activity identity on worker queue entries:
-- Add columns to worker_queue table
ALTER TABLE worker_queue ADD COLUMN instance_id TEXT;
ALTER TABLE worker_queue ADD COLUMN execution_id INTEGER;
ALTER TABLE worker_queue ADD COLUMN activity_id INTEGER;
-- Add index for efficient cancellation
CREATE INDEX idx_worker_queue_activity ON worker_queue(instance_id, execution_id, activity_id);- Implement batch deletion in
ack_orchestration_item:
// Delete cancelled activities atomically within the ack transaction
for activity in cancelled_activities {
DELETE FROM worker_queue
WHERE instance_id = activity.instance
AND execution_id = activity.execution_id
AND activity_id = activity.activity_id;
}Release: https://crates.io/crates/duroxide/0.1.7
-
Cooperative activity cancellation - Activities can detect when their parent orchestration has been cancelled or completed
ActivityContextnow provides cancellation awareness viais_cancelled()andcancelled()methods- Activities can cooperatively respond to cancellation by checking the cancellation token
- Use
tokio::select!withctx.cancelled()for responsive cancellation in async activities - Configurable grace period before forced activity termination
-
ExecutionState enum - Providers now report orchestration state with activity work items
ExecutionState::Running- Orchestration is active, activity should proceedExecutionState::Terminal { status }- Orchestration completed/failed/continued, activity result won't be observedExecutionState::Missing- Orchestration instance deleted, activity should abort
-
Provider validation tests for cancellation - 13 new tests in
provider_validation::cancellation- Verifies
ExecutionStateis correctly returned byfetch_work_itemandrenew_work_item_lock - Tests for Running, Terminal (Completed/Failed/ContinuedAsNew), and Missing states
- Tests for state transitions during activity execution
- Verifies
-
Single-threaded runtime support - Full compatibility with
tokio::runtime::Builder::new_current_thread()- Essential for embedding in single-threaded environments (e.g., pgrx PostgreSQL extensions)
- New scenario tests in
tests/scenarios/single_thread.rs - Use
RuntimeOptions { orchestration_concurrency: 1, worker_concurrency: 1, .. }for 1x1 mode
-
Configurable wait timeout for stress tests -
StressTestConfig::wait_timeout_secsfield- Default: 60 seconds
- Increase for high-latency remote database providers
- Uses
#[serde(default)]for backward compatibility with existing configs
-
BREAKING:
Provider::fetch_work_itemnow returns 4-tuple:(WorkItem, String, u32, ExecutionState)- Added
ExecutionStateas fourth element to report parent orchestration state - Required for activity cancellation support
- Added
-
BREAKING:
Provider::renew_work_item_locknow returnsExecutionStateinstead of()- Allows runtime to detect orchestration state changes during long-running activities
- Triggers cancellation token when orchestration becomes terminal
-
Provider validation test count increased from 62 to 75
-
Documentation updates:
- Added "Runtime Polling Configuration" section to provider-implementation-guide
- Default polling interval (10ms) is aggressive; configure for remote/cloud providers
- Updated provider-testing-guide with new test count and wait_timeout_secs examples
-
test_worker_lock_renewal_extends_timeout - Fixed timing sensitivity (GitHub #34)
- Test now creates proper orchestration with Running status before testing renewal
- Uses 0.6x pre-renewal wait + 0.4x post-renewal wait for reliable timing
-
test_multi_threaded_lock_expiration_recovery - Fixed race condition (GitHub #32)
- Uses
tokio::sync::Barrierto synchronize thread start times - Eliminates false failures from connection pool cold-start latency
- Uses
Provider implementers:
// fetch_work_item now returns ExecutionState
async fn fetch_work_item(
&self,
lock_timeout: Duration,
poll_timeout: Duration,
) -> Result<Option<(WorkItem, String, u32, ExecutionState)>, ProviderError>;
// renew_work_item_lock now returns ExecutionState
async fn renew_work_item_lock(
&self,
token: &str,
extend_for: Duration,
) -> Result<ExecutionState, ProviderError>;Determining ExecutionState:
// Query the execution status for the work item's instance/execution_id
let state = match (instance_exists, execution_status) {
(false, _) => ExecutionState::Missing,
(true, None) => ExecutionState::Missing,
(true, Some(status)) if status == "Running" => ExecutionState::Running,
(true, Some(status)) => ExecutionState::Terminal { status },
};Activity authors (using cancellation):
activities.register("LongTask", |ctx: ActivityContext, input: String| async move {
for item in items {
// Check cancellation periodically
if ctx.is_cancelled() {
return Err("Cancelled".into());
}
process(item).await;
}
Ok("done".into())
});
// Or use select! for responsive cancellation
activities.register("AsyncTask", |ctx: ActivityContext, input: String| async move {
tokio::select! {
result = do_work(input) => result,
_ = ctx.cancelled() => Err("Cancelled".into()),
}
});Release: https://crates.io/crates/duroxide/0.1.6
-
Large payload stress test - New memory-intensive stress test scenario
- Tests large event payloads (10KB, 50KB, 100KB) and longer histories (~80-100 events)
- New binary:
large-payload-stressfor running the test standalone - Uses the same
ProviderStressFactorytrait as parallel orchestrations test - Configurable payload sizes and activity/sub-orchestration counts
- See
docs/provider-testing-guide.mdfor usage
-
Stress test monitoring - Resource usage tracking in
run-stress-tests.sh- Peak RSS (Resident Set Size) measurement
- Average CPU usage tracking
- Sampling every 500ms during test execution
- New documentation:
STRESS_TEST_MONITORING.md - Supports
--parallel-onlyand--large-payloadflags
-
Memory optimization - Reduced allocations in history processing
- Added
HistoryManager::full_history_len()- get count without allocation - Added
HistoryManager::is_full_history_empty()- check emptiness without allocation - Added
HistoryManager::full_history_iter()- iterate without allocation - Updated runtime to use efficient methods in hot paths
- Improved child cancellation to use iterator instead of collecting full history
- Added
-
Orchestration naming - Renamed "FanoutWorkflow" to "FanoutOrchestration" for consistency
- Child sub-orchestration cancellation now uses iterator-based approach for better memory efficiency
Release: https://crates.io/crates/duroxide/0.1.5
-
Provider identity API - Providers now expose
name()andversion()methodsProvider::name()returns provider name (e.g., "sqlite")Provider::version()returns provider version- Default implementations return "unknown" and "0.0.0"
- SQLite provider returns "sqlite" and the crate version
-
Runtime startup banner - Version information logged on startup
- Logs duroxide version and provider name/version
- Example:
duroxide runtime (0.1.4) starting with provider sqlite (0.1.4)
-
Worker queue visibility control - Worker queue now uses
visible_atfor delayed visibility- Added
visible_atcolumn to worker_queue (matches orchestrator queue pattern) abandon_work_itemwith delay now setsvisible_atinstead of keepinglocked_until- Cleaner semantics:
visible_atcontrols when item becomes visible,locked_untilonly for lock expiry - Migration file included for existing databases
- Added
-
New provider validation tests - 2 additional queue semantics tests
test_worker_item_immediate_visibility- Verify newly enqueued items are immediately visibletest_worker_delayed_visibility_skips_future_items- Verify items with future visible_at are skipped
- Reduced default
dispatcher_long_poll_timeoutfrom 5 minutes to 30 seconds- More responsive shutdown behavior
- Better suited for typical workloads
- Provider validation tests - 4 new tests for abandon and poison handling
test_abandon_work_item_releases_lock- Verify abandon_work_item releases lock immediatelytest_abandon_work_item_with_delay- Verify abandon_work_item with delay defers refetchmax_attempt_count_across_message_batch- Verify MAX attempt_count returned for batched messages
- Provider validation test count increased from 58 to 62
abandon_work_itemwith delay now correctly keeps lock_token to prevent immediate refetch
-
Poison message handling - Automatic detection and failure of messages that exceed
max_attempts(default: 10)RuntimeOptions::max_attemptsconfiguration optionErrorDetails::Poisonvariant with detailed contextPoisonMessageTypeenum distinguishing orchestration vs activity poison- Dedicated metrics:
duroxide_orchestration_poison_total,duroxide_activity_poison_total
-
Lock renewal for orchestrations - Prevents lock expiration during long orchestration turns
Provider::renew_orchestration_item_lock()methodRuntimeOptions::orchestrator_lock_renewal_bufferconfiguration (default: 2s)- Automatic background renewal task in orchestration dispatcher
-
Work item abandon with retry - Explicit lock release for failed activities
Provider::abandon_work_item()method with optional delay- Called automatically when
ack_work_itemfails
-
Attempt count management -
ignore_attemptparameter for abandon methodsabandon_work_item(..., ignore_attempt: bool)- decrement count on transient failuresabandon_orchestration_item(..., ignore_attempt: bool)- same for orchestrations- Prevents false poison detection from infrastructure errors
-
Provider validation tests - 8 new poison message tests
orchestration_attempt_count_starts_at_oneorchestration_attempt_count_increments_on_refetchworker_attempt_count_starts_at_oneworker_attempt_count_increments_on_lock_expiryattempt_count_is_per_messageabandon_work_item_ignore_attempt_decrementsabandon_orchestration_item_ignore_attempt_decrementsignore_attempt_never_goes_negative
- BREAKING: SQLite provider is now optional - enable with
features = ["sqlite"] - BREAKING:
Provider::fetch_work_itemnow returns(WorkItem, String, u32)tuple (addedattempt_count) - BREAKING:
Provider::fetch_orchestration_itemnow returns(OrchestrationItem, String, u32)tuple (addedattempt_count) - BREAKING:
Provider::abandon_work_itemnow requiresignore_attempt: boolparameter - BREAKING:
Provider::abandon_orchestration_itemnow requiresignore_attempt: boolparameter OrchestrationItemstruct no longer containslock_token(moved to return tuple)- Provider validation test count increased from 50 to 58
Cargo.toml (if using SQLite provider):
# Before
duroxide = "0.1.1"
# After - SQLite now requires explicit feature
duroxide = { version = "0.1.2", features = ["sqlite"] }Provider implementers:
// fetch_work_item now returns attempt_count
async fn fetch_work_item(...) -> Result<Option<(WorkItem, String, u32)>, ProviderError>;
// fetch_orchestration_item now returns attempt_count
async fn fetch_orchestration_item(...) -> Result<Option<(OrchestrationItem, String, u32)>, ProviderError>;
// abandon methods now have ignore_attempt parameter
async fn abandon_work_item(&self, token: &str, delay: Option<Duration>, ignore_attempt: bool) -> Result<(), ProviderError>;
async fn abandon_orchestration_item(&self, token: &str, delay: Option<Duration>, ignore_attempt: bool) -> Result<(), ProviderError>;
// New method for orchestration lock renewal
async fn renew_orchestration_item_lock(&self, token: &str, extend_for: Duration) -> Result<(), ProviderError>;Runtime users:
RuntimeOptions {
max_attempts: 10, // NEW - poison threshold
orchestrator_lock_renewal_buffer: Duration::from_secs(2), // NEW
..Default::default()
}- Long polling support - Providers can now block waiting for work, reducing CPU usage and latency
dispatcher_long_poll_timeoutconfiguration option (default: 5 minutes)poll_timeout: Durationparameter toProvider::fetch_orchestration_itemandProvider::fetch_work_item- Long polling validation tests in
duroxide::provider_validations::long_polling
- BREAKING:
Provider::fetch_orchestration_itemnow requirespoll_timeout: Durationparameter - BREAKING:
Provider::fetch_work_itemnow requirespoll_timeout: Durationparameter - BREAKING:
RuntimeOptions::dispatcher_idle_sleeprenamed todispatcher_min_poll_interval - BREAKING:
continue_as_new()now returns an awaitable future (usereturn ctx.continue_as_new(input).await)
Provider implementers:
// Add poll_timeout parameter to both fetch methods
async fn fetch_orchestration_item(
&self,
lock_timeout: Duration,
poll_timeout: Duration, // NEW - ignore for short-polling, block for long-polling
) -> Result<Option<OrchestrationItem>, ProviderError>;Runtime users:
// Rename dispatcher_idle_sleep to dispatcher_min_poll_interval
RuntimeOptions {
dispatcher_min_poll_interval: Duration::from_millis(100),
dispatcher_long_poll_timeout: Duration::from_secs(300), // NEW
..Default::default()
}Orchestration authors using continue_as_new:
// Before: ctx.continue_as_new(input);
// After:
return ctx.continue_as_new(input).await;- Initial release
- Deterministic orchestration execution with replay
- Activity scheduling with automatic retries
- Timer support (create_timer)
- Sub-orchestration support
- External event handling
- Continue-as-new for long-running workflows
- SQLite provider implementation
- OpenTelemetry metrics and structured logging
- Provider validation test suite
- Comprehensive documentation