Isbn: 9783389179833 - generating instrument sounds aligned with video via human body keypoints: a deep learning approach to multimodal audio-visual synthesis (5 risultati)

Perfeziona la tua ricerca

  • Libri (5)

  • Nuovo (5)

a

Fascia di prezzo personalizzata (EUR)

a

  • Condizione: Nuovo

    EUR 25,75

     Spedizione gratuita 
    Spedito in U.S.A.

    Quantità: Più di 20 disponibili

    Condizione: New.

  • Lingua: Inglese

    Editore: GRIN Verlag, 2026

    3389179836 / 9783389179833

    • Brossura

    Da: AHA-BUCH GmbH, Einbeck, GermaniaAHA-BUCH GmbH

    Venditore con 5 stelle
    Contatta il venditore

    Condizione: Nuovo

    EUR 18,95

    EUR 60,37 spedizione 
    Spedito da Germania a U.S.A.

    Quantità: 1 disponibili

    Taschenbuch. Condizione: Neu. Druck auf Anfrage Neuware - Printed after ordering - Research Paper (undergraduate) from the year 2026 in the subject Computer Science, grade: Good, , language: English, abstract: Historical video archives and recordings from the past often suffer from degraded or completely missing audio tracks due to deterioration of storage media, recording limitations of the era, or loss during archival processes. Similarly, silent films and performance documentation may lack synchronized sound entirely. Emerging generative artificial intelligence techniques have demonstrated the potential to reconstruct missing audio content by analyzing visual information alone-a capability particularly valuable for restoring cultural heritage materials and historical performance recordings. However, when applied to complex activities such as musical instrument performance, existing methods have shown limited accuracy in capturing the nuances of sound production. Prior research has established that SpecVQGAN architectures combined with Transformer-based mechanisms can improve video-to-audio generation. This work introduces an enhanced model that augments SpecVQGAN by incorporating human skeletal pose features, specifically designed to elevate the quality of generated musical instrument sounds. Through comprehensive evaluation using both subjective user studies and objective quantitative metrics, we demonstrate that the proposed framework significantly outperforms existing approaches in reconstructing authentic instrumental audio from archival and silent performance videos.

  • Lingua: Inglese

    Editore: GRIN Verlag Feb 2026, 2026

    3389179836 / 9783389179833

    • Brossura
    • Print on Demand

    Da: BuchWeltWeit Ludwig Meier e.K., Bergisch Gladbach, GermaniaBuchWeltWeit Ludwig Meier e.K.

    Venditore con 5 stelle
    Contatta il venditore

    Condizione: Nuovo

    EUR 18,95

    EUR 23,00 spedizione 
    Spedito da Germania a U.S.A.

    Quantità: 2 disponibili

    Taschenbuch. Condizione: Neu. This item is printed on demand - it takes 3-4 days longer - Neuware 40 pp. Englisch.

  • Lingua: Inglese

    Editore: GRIN Verlag Feb 2026, 2026

    3389179836 / 9783389179833

    • Brossura
    • Print on Demand

    Da: buchversandmimpf2000, Emtmannsberg, BAYE, Germaniabuchversandmimpf2000

    Venditore con 5 stelle
    Contatta il venditore

    Condizione: Nuovo

    EUR 18,95

    EUR 60,00 spedizione 
    Spedito da Germania a U.S.A.

    Quantità: 1 disponibili

    Taschenbuch. Condizione: Neu. This item is printed on demand - Print on Demand Titel. Neuware -Research Paper (undergraduate) from the year 2026 in the subject Computer Science, grade: Good, , language: English, abstract: Historical video archives and recordings from the past often suffer from degraded or completely missing audio tracks due to deterioration of storage media, recording limitations of the era, or loss during archival processes. Similarly, silent films and performance documentation may lack synchronized sound entirely. Emerging generative artificial intelligence techniques have demonstrated the potential to reconstruct missing audio content by analyzing visual information alone-a capability particularly valuable for restoring cultural heritage materials and historical performance recordings. However, when applied to complex activities such as musical instrument performance, existing methods have shown limited accuracy in capturing the nuances of sound production. Prior research has established that SpecVQGAN architectures combined with Transformer-based mechanisms can improve video-to-audio generation. This work introduces an enhanced model that augments SpecVQGAN by incorporating human skeletal pose features, specifically designed to elevate the quality of generated musical instrument sounds. Through comprehensive evaluation using both subjective user studies and objective quantitative metrics, we demonstrate that the proposed framework significantly outperforms existing approaches in reconstructing authentic instrumental audio from archival and silent performance videos. 40 pp. Englisch.

  • Lingua: Inglese

    Editore: GRIN Verlag, 2026

    3389179836 / 9783389179833

    • Brossura
    • Print on Demand

    Da: preigu, Osnabrück, Germaniapreigu

    Venditore con 5 stelle
    Contatta il venditore

    Condizione: Nuovo

    EUR 18,95

    EUR 70,00 spedizione 
    Spedito da Germania a U.S.A.

    Quantità: 5 disponibili

    Taschenbuch. Condizione: Neu. Generating Instrument Sounds Aligned with Video via Human Body Keypoints | A Deep Learning Approach to Multimodal Audio-Visual Synthesis | Haruka Okano (u. a.) | Taschenbuch | Englisch | 2026 | GRIN Verlag | EAN 9783389179833 | Verantwortliche Person für die EU: preigu GmbH & Co. KG, Lengericher Landstr. 19, 49078 Osnabrück, mail[at]preigu[dot]de | Anbieter: preigu Print on Demand.