Generating Instrument Sounds Aligned with Video via Human Body Keypoints : A Deep Learning Approach to Multimodal Audio-Visual Synthesis

Lingua: inglese

Editore: GRIN Verlag, 2026

3389179836 / 9783389179833

Da: AHA-BUCH GmbH, Einbeck, GermaniaAHA-BUCH GmbH

Venditore con 5 stelle

Venditore AbeBooks dal 14 agosto 2006

Visualizza gli articoli di questo venditore
Brossura

Condizione: Nuovo

EUR 18,95

EUR 60,37 spedizione 
Spedito da Germania a U.S.A.

Quantità: 1 disponibili

Aggiungi al carrello
Resi gratuiti per 30 giorni

Descrizione dell’articolo da parte del venditore

Druck auf Anfrage Neuware - Printed after ordering - Research Paper (undergraduate) from the year 2026 in the subject Computer Science, grade: Good, , language: English, abstract: Historical video archives and recordings from the past often suffer from degraded or completely missing audio tracks due to deterioration of storage media, recording limitations of the era, or loss during archival processes. Similarly, silent films and performance documentation may lack synchronized sound entirely. Emerging generative artificial intelligence techniques have demonstrated the potential to reconstruct missing audio content by analyzing visual information alone-a capability particularly valuable for restoring cultural heritage materials and historical performance recordings. However, when applied to complex activities such as musical instrument performance, existing methods have shown limited accuracy in capturing the nuances of sound production. Prior research has established that SpecVQGAN architectures combined with Transformer-based mechanisms can improve video-to-audio generation. This work introduces an enhanced model that augments SpecVQGAN by incorporating human skeletal pose features, specifically designed to elevate the quality of generated musical instrument sounds. Through comprehensive evaluation using both subjective user studies and objective quantitative metrics, we demonstrate that the proposed framework significantly outperforms existing approaches in reconstructing authentic instrumental audio from archival and silent performance videos.

Codice articolo 9783389179833

Titolo
Generating Instrument Sounds Aligned with Video via Human Body Keypoints : A Deep Learning Approach to Multimodal Audio-Visual Synthesis
Autore
Yasuyuki Tahara
Editore
GRIN Verlag
Anno di pubblicazione
2026
Condizione
Neu
Rilegatura
Taschenbuch
Lingua
inglese
ISBN 10
3389179836
ISBN 13
9783389179833
Peso dell'articolo
73 grammi
Dimensioni
210x148x4 mm

AHA-BUCH GmbH

Einbeck, Germania

Venditore con 5 stelle

Venditore AbeBooks dal 14 agosto 2006

Tariffe di spedizione da Germania a U.S.A.

ArticoloDa 30 a 40 giorni lavorativiDa 7 a 14 giorni lavorativi
Primo articoloEUR 60,37EUR 70,37
I tempi di consegna sono stabiliti dai venditori e variano in base al corriere e al paese. Gli ordini che devono attraversare una dogana possono subire ritardi e spetta agli acquirenti pagare eventuali tariffe o dazi associati. I venditori possono contattarti in merito ad addebiti aggiuntivi dovuti a eventuali maggiorazioni dei costi di spedizione dei tuoi articoli.

Metodi di pagamento

  • Visa
  • Mastercard
  • American Express
  • Carte Bleue
  • Apple Pay
  • Google Pay
  • Assegno
  • Bonifico bancario
  • PayPal

Descrizione dello Store

Das Unternehmen AHA-BUCH GmbH: Seit der Gründung von AHA-BUCH im Juli 2005 ist unser Hauptziel, zufriedenen Kunden so schnell und so preisgünstig wie möglich ihren Bücherwunsch zu erfüllen. Unsere Firma beschäftigt 16 Mitarbeiter, die nur ein Ziel kennen: den Kunden und seine Wünsche! Auf über 3700 m2 Fläche haben wir über 100.000 Bücher, Modernes Antiquariat und Spiele auf Lager.

Specializzazione

Kinderbücher & Kinderhör Casetten, German Books, Software, Natur & Tiere, Ratgeber, Sachbücher, Englische Bücher, Medizin & Gesundheit, Universität & Studium

Informazioni sull’azienda del venditore

AHA-BUCH GmbH

Garlebsen 48
Einbeck, Germania 37574