framework for fusing continuous audio embeddings into a causal language model for audio understanding