Modalities

This page summarizes what is currently shipped in hama (v1.7.0) and what remains on the roadmap.

Available

  • Text → IPA (G2P): Available in Python, Node/Bun, and the browser. Returns IPA plus per-phoneme alignment metadata.

  • Audio → Phoneme (ASR): Available in Python, Node/Bun, and the browser. Accepts waveform input and returns collapsed phoneme output from the asr_waveform.hama model. The current shipped model is the v3 four-block Conformer introduced in v1.7.0, and the default decode temperature is 0.95 with blank bias -0.1. You can also request per-phoneme time spans (phoneme_spans / phonemeSpans) that report each phoneme’s approximate start/end in milliseconds and frames. Spans are coarse because CTC alignment is peaky.

  • IPA → Text (P2G): Available in Python, Node/Bun, and the browser. Phoneme-to-grapheme conversion, the inverse of G2P. The result includes alignments mapping each output token back to the input phoneme it most attends to.

  • Audio → Text: Available in Python, Node/Bun, and the browser. Compose phoneme ASR and P2G using the ASR result’s word-boundary-preserving text; see the Audio → Text example. There is no separate audio-to-text model API.

Coming soon

  • Text → Embeddings

Runtime coverage

RuntimeG2PPhoneme ASRP2G
BrowserYesYesYes
Node/BunYesYesYes
PythonYesYesYes