PhDing ;-)

woensdag, april 02, 2008

The Need For Real Data

Yes, I am working in academic setting. But, no, my work should not only work in academic settings.
I think one of the biggests differences of information retrieval research (at least for a web/worldwide setting) to other disciplines is that big companies have an inherent advantage. We are playing with 50 queries on some hundret thousand documents or even only few ten thousand shots in video retrieval. The relevance of queries get assessed by a single sociological and professional group often situated in the same building. A user study consists of an afternoon of baught (sorry motivated) students which click on some items they are only partially interested.

Simulation!?

dinsdag, maart 25, 2008

What to consider doing Experiments

I find a lot of time is wasted for me, while doing experiments. Often it is small stuff like different line endings, number format, hard to sort things etc. Sometimes it is also more important like different keys for the same thing (upper-/lowercase, hyphens etc). This can really take time to get sorted out. Moreover, it is most likely difficult to reproduce. Therefore, I should unify as soon as possible...

dinsdag, februari 12, 2008

Things to remember when writing a paper

Writing:
  • Think more about the structure first.
  • Latex Macros for common terms don't help much
  • Describe Figures, so that they can be looked at without looking into the Text.
  • Be more verbose than you think
  • Give commonly used mathematical terms a name to refer to - don't repeat the formula
Distributed Writing:
  • Use Subversion to control text even if the others don't know how to use it.
  • Use one paragraph one line format. This captures the semantics what you are doing the best.
  • Make it visible in on the paper which revision the text belonged to
  • latexdiff is a good tool to find out differences about submissions. Take care of people who don't know what a linebreak or a carrige return are.
  • Spellcheck before every commit
  • Chapters (big document parts) in different files. (good for collaboration)
Experiments:
  • Data Organization
    • Keep data and derived data (for the experiments) appart
    • Keep task data (preprocessed queries) apart from task results (rankings)
    • Keep all scripts in one place
  • Visualize the Data Flow
  • Note EVERYTHING down
  • Think first about what you want to show
  • Then choose the parameters you want to display
    • Pay attention to parameters in which you are not interested, but which are still freely choosable.
  • Choose way of display (table, Histogram, Line / Point Plot)
  • Perform the experiment first with a small but representative Dataset to see approximate Results
  • Perform full experiments
  • Check experiments (results for all queries etc)
  • Create Graphs
  •  
Submission preparation (1 Week before)
  • Create account in yet another conference tool.
  • Upload abstract and affiliation

Check before submission:
  • Anonymousness: Check own references and mentioning in text.
  • Spell Check (convert to rtf and then in word)
  • Conference requirements met (Template etc)
  • Graphs have X and Y axis labeled
  • Figure / Table captions coherent state.
  • Hanging Lines (Lines or Headings being isolated from the rest on the end or start of a page / column)
  • Check Reference for completeness
  • Remove Subversion comments
  • Check Coherence of enumeration styles (1) 2) Third...
  • Check all identifiers ($n$) etc are described
  • Sections are listed in introduction & Subsections at the beginning of each section (not necessarily)
  • Title spelled correctly (double check!)
Camera Ready Version
  • Authors / Affiliation / Title spelled correctly (double check!)
  • Copy Bibliography / external tex files into paper folder (so that they don't get changed anymore). Alternatively, paste .bbl file into .text
Cleaning Phase:
  • Delete Unnecessary Data
  • If time permits, also delete intermediate data and redo the whole experiment
  • Backup the whole Experiment Folder (.tex, plots, sources, data, scripts, binaries...)
    You might run the experiment on newer hardware / software for your thesis, but there might be changes which screw your results!

Adjustment of Aims

Future Aims in my PHD:
  • Short Term (Till Summer)
    • Assesment of Detectors (Improve through co-occurrence)
    • Improve P(Ci|R)
    • Integration of ASR output (Laurens, Marijn and me)
    • Work on Database (Ander: uncertain Data) (CIKM)
    • Platform development
  • Mid Term (Till Autumn)
    • Evaluation Buisiness (SIGIR)
    • Internship Abroad (Yahoo / Google / Microsoft / CMU / City Uni Dublin)
      Better late Winter 2008. Take action Summer
  • Teaching (parallel)
    • Master Thesis Supervisor
    • Supervision (Slaves?) (Browsing Tool)

Thesis Title: Information Retrieval on Uncertain Data

vrijdag, augustus 17, 2007

... long time

A long time has passed. Fortunately I have accomplished some aims mentioned before ... But it seems to go slow as hell. Well I will start documenting again:

This first week after the SIGIR and Work/Holiday at home period at home showed to be half productive. As everybody seems to agree is that I need more knowledge in the IR literature. Therefore I started a paper a day motto - which hold one week. I am trying to bring it back to live but it is hard.

The work for TRECVID is also going slowly. The fact that I am the only one developing on the "bigger thing" I came up with seems to cause problems. There was also a major problem: Of course XQuery (where the data should be stored in) doesn't understand mpeg7 timestamps and therefore can't provide > etc relationships... The last thing I noticed today on friday is that MonetDB seems to have some problem with Namespaces....

dinsdag, oktober 17, 2006

Aims

As discussed with Djoerd and Roeland I defined my short-, mid- and longterm goals:

Shortterm Goals
  1. Gt hands on multiple annotation scoringe
  2. Write survey paper (what kind of annotation / detectors do we have)
  3. More background in (MM-) IR retrieval techniques
    1. Get my hands dirty (see above)
    2. Read more stuff
      1. Basic IR stuff
      2. Language Models / Probabilistic Methods
      3. Ranking functions
  4. Background on building a test corpus
  5. Fostering mathmatical and statistical background
  6. Better define my research questions (Poster, etc and formally)

Midterm Goals
  1. Come up with some ideas on how to solve some of the research problems
  2. Write my first paper about it (maybe not so good conference)
  3. Improve my teaching style / learn how to direct people (MSc)
  4. Sommer school on a cool topic ;-)
  5. Create a network of research-fellows
Longterm Goals
  1. Write many good papers oriented after researchgoals
  2. Hava stay abroad (Inspiration for thesis)
  3. Write PhD thesis

First Workshop

After the very troublesome journey the workshop began next day later than we thought (I could have used some sleep ;-) The first afternoon proceeded with the presentation of the test corpora and the creation of topics supervised by the according track coordinator. The workshop specific terms session, track, campaign, task and breakout session aren’t clear to me. The most interesting part of this day were the coffee breaks for me. That is mainly due to networking. I talked mainly to PhD students about their work. I also met a French men who is working in the very same institute as Lux (a good friend of mine in Singapore). In the evening we had a very nice reception in one of the fanciest hotels in town. Ironically enough I set next to a black researcher … and he was working in Groningen. The talks were very interesting and the food very good (although they like a lot of meat)

The next day could be considered as the “real” workshop. There were two parallel sessions one topic for each topic (ad hoc, web, geoclef, speech, image, q@a). I visited the imageClef session as I was interested at what the people from I2R (the institute in Singapore) are doing. Unfortunately I didn’t really find any anchor point for a cooperation – but we will see. After lunch I had my first presentation in front of a scientific audience (not considering talks in front of my groups. It was only 8 minutes but it still was very exciting. My main back draw was the continuous “mmmhhms” - I really have to work on. On the way back to town I sat next to a rather famous IR-researcher-women as it seems. The evening brought a invitation of the City Hall of Alicante. The food was very nice and I met a more senior researcher from I2R who was also French and a very nice person. He even knew Lux personally, good friend of mine.

The third day, Friday, consisted of “breakout sessions” on the various areas. I only attended the speech session. The track seemed to be some kind of stuck as there was no real progress made. The MALACH collection will be expanded to contain more data in the check language. I announced that I will evaluate whether there is interest in my groups to combine GeoCLEF with speech recognition techniques… Let us see how that works out.... The workshop was ended with a set of rather boring talks.


The flight back was not spectacular!

woensdag, augustus 09, 2006

First Writing (CLEF Paper)

After it has finally been decided that I am going to CLEF, I have to get my
feed wet with writing. My first approaches (not surprisingly) were not all that great.

My first paragraphes of the Introduction were like this:
--------------------------------------
In the past the Information Retrieval community spent the most efforts
to enable the user to retrieve documents, which were available as text.
Because of a huge amounts of information exists in spoken documents
and other multimedia content it would be very benefitial for the user
if he was supplied with information from this type of content too.
--
Comment: The first sentence is kind of obvious to the audiance. So it could
be left out. The second part is also leading into wrong type of questions that
will be covered in the paper. It is a CLSR track so everybody should believe that
it is benefitial. It is also kind of obvious to other audiance.
--------------------------------------
Obviously there exists an impedence mismatch between the format of
the documents and the way the user enter his query. Because even if
he uses speech input to express his query, it is unlikely that he
actually means the acustic properties (i.e. pronaunciation) of his
input he is looking for. That is why it is a big issue to bridge the
gap of a query existing in a textual query language and the relevant
documents being stored as acustic waves.
--
Comment: These is also a very general statements. They don't get
investigated further in the paper. The imedence mismatch also was
inspired by my database background (RDBMS vs OO-Programming).
I think it does play a role and it is the motivation to use annotations
(we can't query collection of audio files, at least not online). Another
notion which could be deduced from here is the way of how users enter
there information request (query) into the computer. Specially in the
case of rather structured queries ("somewhere in the middle of the film")
this is not addressed so much to my tiny knowledge.
--------------------------------------

On the structure: My approach was generally first to do some blabla about the
current situation of current research (which I moreover have no glue
of). Then I wanted to introduce the "small" problem which we tried to
solve with PFTijah. Section 2 would contain the intro to PFTijah, its
application to the CLEF-CLSR topics and the procedure of simulating
Dutch people who request the information in dutch. Then I would
have conclude that ASR is benefitial together with MK/SUM and
presented the outline of (my) future work.

Roeland changed this to the introduction of the problem with multiple
annotations in the beginning. Afterwards in section 2 PFTijah and its application
to the task. In the conclusion he closed the "ring" in the way that he repeated what remains
open an what not. He also did remove a lot of to vague and off-topic statements
of mine.

Writing hand-craft: A lesson learned while writing was also to fix each
paragraph to a single statement (little topic). Attention must be paid not
to write on verious items in one paragraph. This makes the paper easier to
read as it is presented in clear units. It also helps to assess the correctness
of the logic used in the paper. This is the case because the punch line through
the paper could be much easier followed and analysed. Furthermore one could benchmark
each paragraph against its "heading". If it contains off-topic stuff it could
mean that:
  1. the paragraph should be split into two units of seperate meaning.
    Afterwards the punchline (at least in the section) should be re-assest.
  2. the additional content is just filling ("nice to have"). So delete it.
  3. the presentation of this aspect is actually useful. Thus the heading
    should be changed to union the touched items. The punchline should
    be re-investigated again
It might also be that two paragraphs (not necessarily following each other) should
be combined. This could be because their content is actually belonging to each other.
The punchline need then obviously to be reexaminied.

woensdag, juli 26, 2006

DB Query Research vs. IR Research

Sorry this really turned out to be "novel" instead of summary.

What surprised me a little at the beginning is that IR researchers seem not to care too much about performance. Being the most important type of benchmark agains other approaches in the DB world, I expected this to be true in IR as well. The explanation why it isn't, goes like this:

In DB research you have a query and a (sometimes semi-) structured dataset. You run queries on the data limiting the result to dataitems which qualify for the constraints given in the query. Given a specific dataset and query the outcome must be always deterministic. And can't be questioned by anymbody. That means that once you are sure that the query, you developed, is the semantically correct, it returns the the exact answer on your query (question) from the underlying data base. In that sense it is like the question answering discipline which extracts few possible answers to a (natural) question from a given source of documents (information base). Of course the difference is that question answering only guesses an answer which could be far from being correct/exact. The opinion about the answer being correct is also very subjective.

In IR systems the user formulates a query as well. The system returns a set of information sources (documents) which are likely to contain the information, which the user searched for in his query. The set of sources is most often ranked by the systems estimated likelyness that the source is relevant. Therefore it forms a list of results with the most relevant sources at the beginning. The problem with the presentation of information sources is that the user might find other documents more relevant. Or he would have done the ranking of the relevant documents in a different way.

From these facts stem the different focuses on the evaluation of the performance of a query. As in the DB field the correctness of a query result is assumed the focus lies on quantitiv measures - mainly the response (speed) time of the DBMS. In the IR approach speed also matters. What is much more important though is a reasonably good fulfillment of the users need for information. As this aspect differes from user to user for the same result a qualitative approach is much more important. It is in the first step of utmost importance that the need for information is optimally fullfilled. Only at the seconde step the responsetime of the systems plays a role.

Semistructured data: The increasing importance of semi structured data brought the DB discipline and the IR discipline closer together. This began with the SGML standard but realy took off with the wide usage of XML on the internet. The user can limit the parts of the data set in which he is searching for information by structural constrains. Structure could also be translated with adding context to the originial information (which had no explicit context before). In the case of XML the context is a keyword identifying the semantic of the data.

Multimedia Context: In Multimedia Context this could also be other things than semantics. Looking on video data the context of information might be the border of one shot to another. To generalize this, the context could be various sorts of events. They could be only one point in time (shot switch) or a time span (silent/noise detection).

Research Question: Given all these events/structures forming the context of a (multimedial) information source we want to find relevant (video-)documents to our query in a collection.
  1. The first question arrises when we look at storing the structure/events. The various events could happen interleaved. There could be a shot event during the duration of a noise event, which could be a commentator talking through out multiple shots. That means, we need a intelligent way of storing these events with fast access and low maintenance time.
  2. The next question could be how do we optimally exploit these structures to answer the queries. (i.e. what kind of index-structures should we use and kind of query processing model do we employ.
  3. And as a thrird question we will need to consider whether the available query lanuages (from DB / IR) are adquate to express the queries that fullfil the user's information need.