Similar presentations:
project_proposal_defense_mark_bondarenko
1.
PROJECT PROPOSAL DEFENSEMus ic E diting
Us ing the
P rompt-to-Prompt
Approach
Can a diffusion model change one
musical attribute without
rebuilding the whole track?
Mark Bondarenko
HSE University
1
2.
Why this problem mattersMOTIVATION
Generation quality is improving
fast.
Precise editing is still the hard
part.
Edi t i ng mus t pr es er v e t empo,
ar r angement , and over al l c oher enc e.
The des i r ed change i s of t en l oc al :
l y r i c s , i ns t r ument at i on, s t y l e, or
one component .
Nai v e r e-gener at i on t ends t o
modi f y muc h mor e t han i nt ended.
From-scratch generation = freedom
Editing an existing track = precision +
preservation
Tar get out c ome: l oc al edi t ,
gl obal s t r uc t ur e i nt ac t
2
3.
F rom image editing to mus ic editingRELATED WORK AND GAP
Prompt-to-Prompt
(vision)
Tr ai ni ng-f r ee
semant i c edi t s
t hr ough cr oss at t ent i on cont r ol
dur i ng denoi si ng.
PPAE
(audio)
Inver t -and-edi t
pi pel i ne
f or l oc al i zed audi o
changes
wi t h at t ent i on-l ev el
i nt er vent i on.
Open gap
M
oder n t ext -t o-musi c
(music)
backbones
have not been
syst emat i cal l y
t est ed
f or pr eci se l ocal
musi c edi t s.
Research gap: can Prompt-to-Prompt-style cross-attention control work on structured
musical material using a mo dern backbone such as ACE-Step v1.5?
3
4.
R es earch ques tion and core ideaWHAT THE PR OJ ECT TES TS
Research question
Can cross-attention
injection enable localized
text-guided mus ic edits
while preserving the
global mus ical form?
Source prompt
"s y nt h-pop
t r ac k
wi t h
f emal e
voc al s "
Target prompt
Attention maps
Edited output
"s ame t r ac k
wi t h s of t er
dr ums "
Rec or d t ok enl ev el
condi t i oni ng
dur i ng
denoi s i ng
Loc al
change
non-t ar get
cont ent
s t abl e
Working hypothesis
Weak er i nt er v ent i on at
ear l i er , s t r uc t ur e-l ev el
s t ages and s t r onger
i nt er v ent i on l at er may
impr ov e l oc al edi t abi l i t y
wi t hout br eak i ng mel ody ,
t imi ng, or c oher enc e.
Injec t or r epl ac e
s el ec t ed maps
at chos en
l ay er s and t ime
s t eps
4
5.
Why ACE -S tep v1 .5 — and why this isfeas
ible now
BACKBONE AND READINESS
Why ACE-Step v1.5
1
2
3
Moder n open t ext -t o-musi c
bac kbone wi t h st r ong
gener at i on qual i t y.
Posi t i oned as a br oader musi c
mani pul at i on f r amewor k, not
onl y pur e gener at i on.
A r eal i st i c t est bed f or
cont r ol l abl e, pr ompt -dr i ven
edi t i ng.
Prototype already
available
ACE-St ep v1.0 → gener at i on, Nul l -Text
i nv er s i on, i ni t i al
edi t i ng
r epl ac ement -bas ed
Immediate next step: migrate the prototype to
ACE-Step v1.5 and expose cross-attention hooks
for systematic experiments.
5
6.
Methodology in one pipelineHOW THE S TUDY IS ORGANIZED
1
2
3
4
5
Integrate
editing
hooks
Expos
e cr os s -
Invert real
audio
Record
attention
maps
Control
denoising
Compare
outputs
at t ent i on
ac t i v at i ons
i n ACE-St ep
v1 .5.
Variables under study
Map i nput
mus i c t o t he
l at ent /noi s e
t r ajec t or y .
attention layer choice
Tr ac k s our c epr ompt t ok en
condi t i oni ng.
token mapping
Injec t /
r epl ac e
s el ec t ed maps
by l ay er and
s t ep.
denoising-step schedule
Tes t
s y nt het i c and
r eal audi o
under
mul t i pl e
var i ant s .
synthetic vs real audio
6
7.
E x perimental des ignWHAT WILL BE TESTED AND HOW
Edit families
Baselines
Evaluation
Nai ve pr ompt
r epl acement
Edi t success
St yl e / genr e / t i mbr e
shi f t s
Var i ant wi t hout
i nver si on
Pr eser vat i on of musi cal
st r uct ur e
Component -l evel changes:
vocal s, dr ums, bass,
gui t ar
Var i ant wi t h r est r i ct ed
at t ent i on mani pul at i on
Per cept ual coher ence of
t he edi t ed r esul t
Wher e possi bl e:
per f ormer or vocal
t i mbr e changes
Component -wi se
compar i son t o i sol at e
each cont r i but i on
Compar i son on synt het i c
t r acks and r eal audi o
Lyr i c -l evel
edi t s
7
8.
E x pected contribution — with clear ris ksWHAT A SUCCES SFUL PR OJ ECT DELIVER S
Expected contribution
Main risks / limitations
A r epr oduc i bl e t r ai ni ng-f r ee edi t i ng
pi pel i ne f or ACE-St ep v1 .5.
Met hods t r ans f er r ed f r om images may not
map c l eanl y t o mus i c .
Ev i denc e on wh en Pr ompt -t o-Pr ompt s t y l e cont r ol wor k s bes t i n mus i c .
Inv er s i on c an i nt r oduc e ar t i f ac t s on
r eal audi o.
A c l ear er vi ew of t he t r ade-of f
bet we en edi t s t r engt h and s t r uc t ur e
pr es er v at i on.
Cur at ed exampl es t hat s how bot h
s uc c es s f ul edi t s and cur r ent
l imi t at i ons .
Loc al edi t s may s t i l l di s t ur b gl obal
ar r angement or coher enc e.
Fi nal qu al i t y r emai ns cons t r ai ned by
t he bac k bone model i t s el f .
A strong defense message is not “the method will definitely
solve everything”, but “the project is well-posed, testable,
and likely to produce useful evidence.”
8
9.
TakeawayThis project tests whether
music can be edited
as precisely as images — by
controlling
cross-attention instead of
retraining the model.
If s uc c es s f ul , t he cont r i but i on i s
a cont r ol l abl e, t r ai ni ng-f r ee,
mus i c al l y coher ent edi t i ng
pi pel i ne on t op of a moder n t ext t o-mus i c bac k bone.
Thank you
9